OpenClaw Kubernetes Operator
Self-host OpenClaw AI agents on Kubernetes with production-grade security, observability, and lifecycle management.
OpenClaw is an AI agent platform that acts on your behalf across Telegram, Discord, WhatsApp, and Signal. It manages your inbox, calendar, smart home, and more through 50+ integrations. While Paperclip Inc. offers fully managed hosting, this operator lets you run OpenClaw on your own infrastructure with the same operational rigor.
Why an Operator?
Deploying AI agents to Kubernetes involves more than a Deployment and a Service. You need network isolation, secret management, persistent storage, health monitoring, optional browser automation, and config rollouts, all wired correctly. This operator encodes those concerns into a single OpenClawInstance custom resource so you can go from zero to production in minutes:
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawInstance
metadata:
name: my-agent
spec:
envFrom:
- secretRef:
name: openclaw-api-keys
storage:
persistence:
enabled: true
size: 10Gi
The operator reconciles this into a fully managed stack of 9+ Kubernetes resources: secured, monitored, and self-healing.
Agents That Adapt Themselves
Agents can autonomously install skills, patch their config, add environment variables, and seed workspace files - all through the Kubernetes API, validated by the operator on every request.
# 1. Enable self-configure on the instance
spec:
selfConfigure:
enabled: true
allowedActions: [skills, config, envVars, workspaceFiles]
# 2. The agent creates this to install a skill at runtime
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawSelfConfig
metadata:
name: add-fetch-skill
spec:
instanceRef: my-agent
addSkills:
- "@anthropic/mcp-server-fetch"
Every request is validated against the instance's allowlist policy. Protected config keys cannot be overwritten, and denied requests are logged with a reason. See Self-configure for details.
Note: WithoutselfConfigureenabled, config or skill changes made by the agent inside the container won't trigger a pod restart. You'll need to restart the pod manually (e.g.kubectl delete pod) for changes to take effect.
Features
| | Feature | Details |
|---|---|---|
| Declarative | Single CRD | One resource defines the entire stack: StatefulSet, Service, RBAC, NetworkPolicy, PVC, PDB, Ingress, and more |
| Adaptive | Agent self-configure | Agents autonomously install skills, patch config, and adapt their environment via the K8s API - every change validated against an allowlist policy |
| Secure | Hardened by default | Non-root (UID 1000), read-only root filesystem, all capabilities dropped, seccomp RuntimeDefault, default-deny NetworkPolicy, validating webhook |
| Observable | Built-in metrics | Prometheus metrics, ServiceMonitor integration, structured JSON logging, Kubernetes events |
| Flexible | Provider-agnostic config | Use any AI provider (Anthropic, OpenAI, or others) via environment variables and inline or external config |
| Config Modes | Merge or overwrite | overwrite replaces config on restart; merge deep-merges with PVC config, preserving runtime changes. Config is restored on every container restart via init container. |
| Force Paths | Operator-owned paths under merge | config.forcePaths lists dot-paths the init container rebuilds from the CR on every restart even under mergeMode: merge -- lets managed deployers keep operator-owned config (auth, allowed providers, sandbox image) immune to tenant edits while user-owned config persists |
| Skills | Declarative install | Install ClawHub skills, npm packages, or GitHub-hosted skill packs via spec.skills - supports npm: and pack: prefixes. Additional workspaces can declare their own workspace-scoped skills via additionalWorkspaces[].skills |
| Plugins | Declarative install | Install OpenClaw plugins via spec.plugins - resolved through the OpenClaw CLI ClawHub installer in a secure init container |
| Runtime Deps | pnpm & Python/uv | Built-in init containers install pnpm (via corepack) or Python 3.12 + uv for MCP servers and skills |
| Auto-Update | OCI registry polling | Opt-in version tracking: checks the registry for new semver releases, backs up first, rolls out, and auto-rolls back if the new version fails health checks |
| Scalable | Auto-scaling | HPA integration with CPU and memory metrics, min/max replica bounds, automatic StatefulSet replica management |
| Operational | Instance suspension | Scale to zero with spec.suspended: true - all non-runtime resources remain managed, resume instantly with false |
| Resilient | Self-healing lifecycle | PodDisruptionBudgets, health probes, automatic config rollouts via content hashing, 5-minute drift detection |
| Disk-Aware Readiness | Opt-in ENOSPC guard | spec.probes.diskReadiness renders the readiness probe as an exec check that ANDs the gateway /readyz signal with a workspace writability + free-space check, so a full or read-only PVC drains the pod from Service endpoints instead of silently accepting writes it cannot persist. Liveness/startup stay HTTP so a full disk never turns into a CrashLoopBackOff. Defaulted off. |
| Backup/Restore | S3-backed snapshots | Automatic backup to S3-compatible storage on deletion, pre-update, and on a cron schedule; restore into a new instance from any snapshot |
| Workspace Seeding | Initial files & dirs | Pre-populate the workspace with files and directories before the agent starts; reference an external ConfigMap for GitOps workflows |
| Gateway Auth | Auto-generated tokens | Automatic shared-secret gateway authentication with a persistent token Secret per instance |
| Tailscale | Tailnet access | Expose via Tailscale Serve or Funnel with SSO auth - no Ingress needed |
| Extensible | Sidecars & init containers | Chromium for browser automation, Ollama for local LLMs, Tailscale for tailnet access, plus custom init containers and sidecars |
| Cloud Native | SA annotations & CA bundles | AWS IRSA / GCP Workload Identity via ServiceAccount annotations; CA bundle injection for corporate proxies |
| Cluster Defaults | Singleton CR | OpenClawClusterDefaults (name cluster) fills in unset instance fields - ideal for air-gapped / China regions where every instance would otherwise duplicate the same registry + mirror env boilerplate. Per-instance fields always win. |
| Zombie Reaping | Shared PID namespace | spec.shareProcessNamespace defaults to true so the pause container becomes PID 1 and reaps defunct helper processes from QMD, git, plugins, and shells - no custom init image needed |
Architecture
+-----------------------------------------------------------------+
| OpenClawInstance CR OpenClawSelfConfig CR |
| (your declarative config) (agent self-modification requests) |
+---------------+-------------------------------------------------+
| watch
v
+-----------------------------------------------------------------+
| OpenClaw Operator |
| +-----------+ +-------------+ +----------------------------+ |
| | Reconciler| | Webhooks | | Prometheus Metrics | |
| | | | (validate | | (reconcile count, | |
| | creates -> | & default)| | duration, phases) | |
| +-----------+ +-------------+ +----------------------------+ |
+---------------+-------------------------------------------------+
| manages
v
+-----------------------------------------------------------------+
| Managed Resources (per instance) |
| |
| ServiceAccount -> Role -> RoleBinding NetworkPolicy |
| ConfigMap PVC PDB ServiceMonitor |
| GatewayToken Secret |
| |
| StatefulSet |
| +-----------------------------------------------------------+ |
| | Init: config -> pnpm -> python -> skills* -> custom | |
| | (* = opt-in) | |
| +------------------------------------------------------------+ |
| | OpenClaw Container Gateway Proxy (nginx) | |
| | Chromium (opt) / Ollama (opt) | |
| | Tailscale (opt) + custom sidecars | |
| +------------------------------------------------------------+ |
| |
| Service (default: 18789, 18793 or custom) -> Ingress (opt) |
+-----------------------------------------------------------------+
Quick Start
Prerequisites
- Kubernetes 1.28+
- Helm 3
1. Install the operator
helm install openclaw-operator \
oci://ghcr.io/paperclipinc/charts/openclaw-operator \
--namespace openclaw-operator-system \
--create-namespace
Alternative: install with Kustomize
# Install CRDs
make install
Deploy the operator
make deploy IMG=ghcr.io/paperclipinc/openclaw-operator:latest
Restrict the operator to specific namespaces
To run the operator with namespaced RBAC instead of cluster-wide permissions,
list the namespaces it should watch. The chart switches the namespace-scoped
permissions from a ClusterRole/ClusterRoleBinding to per-namespace
Role/RoleBinding, and passes --watch-namespaces to the operator so its
informer cache is scoped to that list. The operator's own namespace is added to
the Secret informer only, so it can still read its backup credentials, and
the chart renders a matching Secret-only Role there; no other resource type is
watched or granted in the operator namespace. A ClusterRole/ClusterRoleBinding is still created
for the cluster-scoped OpenClawClusterDefaults resource, which the operator
watches regardless of namespace scoping -- a namespaced Role cannot grant
access to a cluster-scoped resource.
helm install openclaw-operator \
oci://ghcr.io/paperclipinc/charts/openclaw-operator \
--namespace openclaw-operator-system \
--create-namespace \
--set 'watchNamespaces={team-a,team-b}'
Each listed namespace must already exist; the chart does not create them.
To bring your own RBAC entirely (e.g. managed by a separate controller or SecurityCenter policy), disable chart-managed RBAC:
helm install openclaw-operator \
oci://ghcr.io/paperclipinc/charts/openclaw-operator \
--namespace openclaw-operator-system \
--create-namespace \
--set rbac.create=false
The kubebuilder markers in internal/controller/ and the manager rules helper
at charts/openclaw-operator/templates/_helpers.tpl document the minimum
permission set the operator requires.
2. Create a secret with your API keys
apiVersion: v1
kind: Secret
metadata:
name: openclaw-api-keys
type: Opaque
stringData:
ANTHROPIC_API_KEY: "sk-ant-..."
3. Deploy an OpenClaw instance
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawInstance
metadata:
name: my-agent
spec:
envFrom:
- secretRef:
name: openclaw-api-keys
storage:
persistence:
enabled: true
size: 10Gi
kubectl apply -f secret.yaml -f openclawinstance.yaml
4. Verify
kubectl get openclawinstances
NAME PHASE AGE
my-agent Running 2m
kubectl get pods
NAME READY STATUS AGE
my-agent-0 1/1 Running 2m
Configuration
Inline config (openclaw.json)
spec:
config:
raw:
agents:
defaults:
model:
primary: "anthropic/claude-sonnet-4-20250514"
sandbox: true
session:
scope: "per-sender"
External ConfigMap reference
spec:
config:
configMapRef:
name: my-openclaw-config
key: openclaw.json
Config changes are detected via SHA-256 hashing and automatically trigger a rolling update. No manual restart needed.
Gateway proxy
By default, each pod includes an nginx reverse proxy sidecar that forwards traffic to the OpenClaw gateway on loopback. Set spec.gateway.enabled: false to disable it:
spec:
gateway:
image:
repository: docker.io/library/nginx
tag: 1.27-alpine
# digest: sha256:... # takes precedence over tag
resources:
requests:
cpu: 10m
memory: 16Mi
limits:
cpu: 100m
memory: 64Mi
The image and resource fields are optional. Omit them to retain the defaults above, or set an image digest to make the proxy supply-chain reference immutable.
- Health probes and Service ports target the gateway directly on port 18789
gateway.bindis set to0.0.0.0instead of loopback- The
gateway-proxycontainer and its tmp volume are omitted from the pod - To replace the built-in proxy with your own (e.g., Envoy, a signing proxy), disable it and add your proxy via
spec.sidecars - Warning: Do not set
gateway.bind: loopbackin your config JSON when the proxy is disabled - the gateway will only listen on127.0.0.1with nothing forwarding external traffic, making the pod unreachable. The operator emits aGatewayBindConflictwarning event if this misconfiguration is detected. - TLS: When the proxy is disabled, the gateway serves plaintext
ws://on0.0.0.0. Ensure your replacement proxy or Ingress handles TLS termination to avoid exposing unencrypted WebSocket traffic (CWE-319).
Disk-aware readiness
By default the readiness probe is an HTTP GET /readyz against the gateway. For PVC-backed instances, /readyz can stay green while the workspace volume is full or read-only (ENOSPC), so the pod keeps receiving traffic while workspace writes fail. Enable the opt-in disk-aware readiness guard to turn the readiness probe into an exec check that combines the gateway /readyz signal with a workspace writability and free-space check:
spec:
probes:
diskReadiness:
enabled: true # default: false (existing deployments are unchanged when unset)
path: /home/openclaw/.openclaw # optional; defaults to the workspace data mount
minFree: 128Mi # optional; minimum free space, a Kubernetes quantity (default 64Mi)
- When enabled, the readiness probe becomes
sh -cexec script that (a) verifiespathis writable (test -w), (b) checks free space withdfagainstminFree, then (c) defers to the gateway/readyzon the same loopback port the HTTP probe would use. The pod is Ready only if all checks pass; the script fails closed (non-zero exit) on any failure. - Liveness and startup stay HTTP-only (
GET /healthz), so a full PVC yieldsNotReady(draining the pod from Service endpoints) rather than a restart loop / CrashLoopBackOff. - The exec script uses only POSIX
sh,test,df, andawk. The/readyzHTTP call usescurlorwgetif present and is skipped gracefully if neither is in the image, so a missing HTTP client never makes a healthy pod permanentlyNotReady(disk checks still run). - This is a secondary, defense-in-depth guard; the application-level
/readyzendpoint remains the primary readiness signal.
Gateway authentication
The operator automatically generates a gateway token Secret for each instance and injects it into both the config JSON (gateway.auth.mode: token) and the OPENCLAW_GATEWAY_TOKEN env var. The token authenticates the gateway connection. A browser's Control UI device identity and one-time pairing are separate security checks.
- The token is generated once and never overwritten - rotate it by editing the Secret directly
- If you set
gateway.auth.tokenin your config orOPENCLAW_GATEWAY_TOKENinspec.env, your value takes precedence - To bring your own token Secret, set
spec.gateway.existingSecret- the operator will use it instead of auto-generating one (the Secret must have a key namedtoken) - The operator sets
OPENCLAW_DISABLE_BONJOUR=1because mDNS discovery is not useful in Kubernetes. This does not disable Control UI device identity. - The operator sets
gateway.mode: local, which current OpenClaw releases require for a gateway that owns local state. Exposure remains controlled by the Service, Ingress, mesh, and gateway bind settings. - Current OpenClaw releases ignore the retired
gateway.controlUi.dangerouslyDisableDeviceAuthsetting, so the operator does not emit it. - On the first browser connection, approve the pending device once from an administrative workstation:
kubectl exec -n <namespace> <instance>-0 -c openclaw -- openclaw devices list
kubectl exec -n <namespace> <instance>-0 -c openclaw -- openclaw devices approve <request-id>
- Supplying the gateway token, including in the Control UI, does not replace browser device pairing.
- Since v2026.2.24, OpenClaw restricts
gateway.allowedOriginsto same-origin by default - if accessing via a non-default hostname (e.g. Ingress), setgateway.allowedOrigins: ["*"]in your config
Control UI allowed origins
The operator auto-injects gateway.controlUi.allowedOrigins so the Control UI works through reverse proxies without CORS errors. Origins are derived from:
- Localhost (always):
http://localhost:18789,http://127.0.0.1:18789for port-forwarding - Ingress hosts: scheme determined from TLS config (
https://if TLS,http://otherwise) - Explicit extras:
spec.gateway.controlUiOriginsfor custom proxy URLs
gateway.controlUi.allowedOrigins directly in your config JSON, the operator will not override it.
Chromium sidecar
Enable headless browser automation for web scraping, screenshots, and browser-based integrations:
spec:
chromium:
enabled: true
image:
repository: chromedp/headless-shell # default
tag: "stable"
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "1000m"
memory: "2Gi"
# Pass extra flags to the Chromium process (appended to built-in anti-bot defaults)
extraArgs:
- "--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
# Inject extra environment variables into the sidecar
extraEnv:
- name: DISPLAY
value: ":99"
When enabled, the operator automatically:
- Injects a
CHROMIUM_URLenvironment variable into the main container - Configures browser profiles in the OpenClaw config - both
"default"and"chrome"profiles are set to point at the sidecar's CDP endpoint, so browser tool calls work regardless of which profile name the LLM passes - Sets up shared memory, security contexts, and health probes for the sidecar
- Applies anti-bot-detection flags by default (
--disable-blink-features=AutomationControlled,--disable-features=AutomationControlled,--no-first-run)
Persistent browser profiles
By default, all browser state (cookies, localStorage, session tokens) is lost on pod restart. Enable persistence to retain browser profiles across restarts:
spec:
chromium:
enabled: true
persistence:
enabled: true # default: false
storageClass: "" # optional - uses cluster default if empty
size: "1Gi" # default: 1Gi
existingClaim: "" # optional - use a pre-existing PVC
When persistence is enabled, the operator creates a dedicated PVC and passes --user-data-dir=/chromium-data to Chrome so that cookies, localStorage, IndexedDB, cached credentials, and session tokens survive pod restarts. This is useful for authenticated browser automation, MFA-protected services, and long-running browser workflows.
Security note: Persistent browser profiles contain sensitive session tokens. The PVC has the same security posture as other instance volumes. Ensure your StorageClass supports encryption at rest for sensitive workloads.
Ollama sidecar
Run local LLMs alongside your agent for private, low-latency inference without external API calls:
spec:
ollama:
enabled: true
models:
- llama3.2
- nomic-embed-text
gpu: 1
storage:
sizeLimit: 30Gi
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "4"
memory: "16Gi"
When enabled, the operator:
- Injects an
OLLAMA_HOSTenvironment variable into the main container - Pre-pulls specified models via an init container before the agent starts
- Configures GPU resource limits when
gpuis set (nvidia.com/gpu) - Mounts a model cache volume (emptyDir by default, or an existing PVC via
storage.existingClaim)
Web terminal sidecar
Provide browser-based shell access to running instances for debugging and inspection without requiring kubectl exec:
spec:
webTerminal:
enabled: true
readOnly: false
credential:
secretRef:
name: my-terminal-creds
resources:
requests:
cpu: "50m"
memory: "64Mi"
limits:
cpu: "200m"
memory: "128Mi"
When enabled, the operator:
- Injects a ttyd sidecar container on port 7681
- Mounts the instance data volume at
/home/openclaw/.openclawso you can inspect config, logs, and data files - Adds the web terminal port to the Service and NetworkPolicy for external access
- Supports basic auth via a Secret with
usernameandpasswordkeys - Supports read-only mode (
readOnly: true) for production environments where shell input should be disabled
Tailscale integration
Expose your instance via Tailscale Serve (tailnet-only) or Funnel (public internet) - no Ingress or LoadBalancer needed:
spec:
tailscale:
enabled: true
mode: serve # "serve" (tailnet only) or "funnel" (public internet)
authKeySecretRef:
name: tailscale-auth
authSSO: true # allow passwordless login for tailnet members
hostname: my-agent # defaults to instance name
image:
repository: ghcr.io/tailscale/tailscale # default
tag: latest
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 200m
memory: 256Mi
When enabled, the operator runs a Tailscale sidecar (tailscaled) that handles serve/funnel declaratively via TS_SERVE_CONFIG. An init container copies the tailscale CLI binary to a shared volume so the main container can call tailscale whois for SSO authentication. The sidecar runs in userspace mode (TS_USERSPACE=true) - no NET_ADMIN capability needed.
State persistence: Tailscale node identity and TLS certificates are automatically persisted to a Kubernetes Secret () via TS_KUBE_SECRET. This prevents hostname incrementing (device-1, device-2, ...) and Let's Encrypt certificate re-issuance across pod restarts. The operator pre-creates the state Secret, grants the pod's ServiceAccount get/update/patch access to it, and mounts the SA token automatically.
Use ephemeral+reusable auth keys from the Tailscale admin console. When authSSO is enabled, tailnet members can authenticate without a gateway token.
NetBird integration
NetBird is a self-hostable alternative to Tailscale: the same WireGuard data plane, with a control plane you can run yourself.
spec:
netbird:
enabled: true
setupKeySecretRef:
name: netbird-setup-key # key: "setupkey"
managementURL: https://netbird.example.com:33073 # omit for NetBird's hosted control plane
hostname: my-agent # defaults to the instance name
Use a reusable, ephemeral setup key from the NetBird dashboard. The operator injects it as NB_SETUP_KEY from the referenced Secret -- it is never written into the pod spec as a literal -- and rolls the pod when the Secret changes.
The sidecar runs in netstack (userspace) mode, so it keeps the same Restricted PSS posture as every other container the operator builds: all capabilities dropped, read-only root filesystem, non-root, seccomp RuntimeDefault. A kernel-mode peer would need NET_ADMIN and /dev/net/tun, which is a different security decision than this operator makes by default.
Peer state lives on an emptyDir, so the peer re-enrolls on restart (which a reusable setup key handles). Unlike Tailscale, NetBird needs no Kubernetes API access, so no ServiceAccount token is mounted and no state Secret is created.
Mesh providers are mutually exclusive. Enabling both tailscale and netbird is rejected by the validating webhook: two overlay clients in one pod would race for the same egress rules and the agent's routing.
Feature comparison:
| | Tailscale | NetBird |
|---|---|---|
| Self-hostable control plane | no | yes (managementURL) |
| Credential | auth key | setup key |
| Serve/Funnel ingress | yes (mode) | not applicable |
| Gateway SSO (authSSO) | yes | no identity header equivalent |
| Needs Kubernetes API | yes (state Secret) | no |
| Persistent node identity | yes (state Secret) | re-enrolls on restart |
Both are implementations of one internal MeshProvider interface, so a third provider means implementing that interface and adding one table entry -- not another copy of the StatefulSet, NetworkPolicy, RBAC and config-enrichment paths.
Config merge mode
By default, the operator overwrites the config file on every pod restart. Set mergeMode: merge to deep-merge operator config with existing PVC config, preserving runtime changes made by the agent:
spec:
config:
mergeMode: merge
raw:
agents:
defaults:
model:
primary: "anthropic/claude-sonnet-4-20250514"
Caveat: In merge mode, removing a key from the CR does not remove it from the PVC config - the old value persists because deep-merge only adds or updates keys. If you need to remove a stale config key, temporarily switch to mergeMode: overwrite, apply, wait for the pod to restart, then switch back to merge.
Partial overwrite under merge mode (forcePaths)
Under mergeMode: merge the operator preserves runtime changes the agent (or a tenant via the Control UI) wrote into the config file. For managed multi-tenant deployments this is a problem: a tenant can persist arbitrary values into operator-owned subtrees -- for example models.providers. -- and route inference through their own third-party key while consuming the deployer's compute.
spec.config.forcePaths is the partial-overwrite escape hatch. For each listed dot-path the init container deletes that subtree from the PVC config and re-applies it from spec.config.raw on every pod restart, so listed paths always match the CR while everything else still persists.
spec:
config:
mergeMode: merge
forcePaths:
- gateway
- models.providers
- agents.defaults.sandbox
raw:
gateway:
auth:
mode: token
models:
providers:
openai:
baseUrl: "https://api.openai.com"
agents:
defaults:
sandbox: true
With the above, channels., settings., and any other user-owned path the agent writes via the Control UI persists across pod restarts; gateway., models.providers., and agents.defaults.sandbox are rebuilt from the CR on every reconcile.
forcePaths is only valid under mergeMode: merge. The validating webhook rejects it under overwrite (where the whole file is already rebuilt every restart) and rejects malformed paths (empty segments, leading or trailing dot, characters outside [a-zA-Z0-9._-]). The same logic runs in both the init container (on pod restart) and the postStart lifecycle hook (on container restart without pod recreation), so an attacker cannot bypass the contract by triggering one form of restart over another.
Skill installation
Install skills declaratively. The operator runs an init container that fetches each skill before the agent starts. Entries use ClawHub by default, or prefix with npm: to install from npmjs.com. ClawHub installs are idempotent - if a skill is already installed (e.g., when using persistent storage), it is skipped rather than failing:
spec:
skills:
- "@anthropic/mcp-server-fetch" # ClawHub (default)
- "npm:@openclaw/matrix" # npm package from npmjs.com
npm lifecycle scripts are disabled globally on the init container (NPM_CONFIG_IGNORE_SCRIPTS=true) to mitigate supply chain attacks.
Skill packs
Skill packs bundle multiple files (SKILL.md, scripts, config) into a single installable unit hosted on GitHub. Use the pack: prefix with owner/repo/path format:
spec:
skills:
- "pack:paperclipinc/skills/image-gen" # latest from default branch
- "pack:paperclipinc/skills/[email protected]" # pinned to tag
- "pack:myorg/private-skills/custom-tool@main" # private repo (requires GITHUB_TOKEN)
Packs are resolved in one of two modes:
1. Manifest mode (explicit) -- the pack path contains a skillpack.json describing which files to seed and where:
{
"files": {
"skills/image-gen/SKILL.md": "SKILL.md",
"skills/image-gen/scripts/generate.py": "scripts/generate.py"
},
"directories": ["skills/image-gen/scripts"],
"config": {
"image-gen": {"enabled": true}
}
}
2. Raw-repo mode (autodiscovery) -- when no skillpack.json is present and the pack path contains a SKILL.md, the operator installs the entire directory verbatim into skills/ in the workspace. This is useful for multi-skill repositories like fluxcd/agent-skills that follow a conventional skills/ layout without per-skill manifests:
spec:
skills:
- "pack:fluxcd/agent-skills/skills/gitops-repo-audit@main"
# installs every file under skills/gitops-repo-audit/ into the workspace
# at skills/gitops-repo-audit/ (including nested assets, schemas, etc.)
Raw mode does not inject config entries into config.raw.skills.entries -- use manifest mode if you need that. The operator refuses to install if GitHub truncates the tree response for very large repositories (add a skillpack.json manifest in that case).
The operator resolves packs via the GitHub Contents + Git Trees APIs (cached for 5 minutes), seeds files into the workspace via the init container, and (in manifest mode) injects config entries into config.raw.skills.entries with user overrides taking precedence. Set GITHUB_TOKEN on the operator deployment for private repo access.
Updating pack contents. By default (spec.skillPackUpdatePolicy: Replace), pack-seeded files converge to the declared pack revision on every pod start: changing a pinned @tag/@commit (or pushing to a tracked branch) overwrites the seeded files, and files that are no longer part of any declared pack are removed. The operator tracks what it seeded in a manifest at /data/.skillpack-manifest on the data volume, so user-created workspace files are never touched. Files at pack-declared paths are operator-managed -- local edits to them are reverted on restart. Set spec.skillPackUpdatePolicy: CreateOnly to opt out and keep the legacy seed-once behavior (files are never overwritten or removed after first seeding; updating a pinned revision then has no effect on already-seeded files).
Workspace-scoped skills (multi-agent). Additional workspaces can declare their own skills with the same reference formats as spec.skills:
spec:
workspace:
additionalWorkspaces:
- name: secondary
skills:
- "pack:example-org/openclaw-skills/skills/[email protected]"
- "@acme/browser-use" # ClawHub, installed into workspace-secondary/skills/
- "npm:@acme/cli-tool" # npm binaries are global (~/.local/bin), shared by all agents
pack: entries resolve exactly like top-level packs (private repos via GITHUB_TOKEN, pinned tags/commits) but seed into ~/.openclaw/workspace- and are tracked in a per-workspace manifest (/data/.skillpack-manifest-ws-), so skillPackUpdatePolicy applies per workspace. ClawHub entries are installed with clawhub --workdir so the skill lands in that workspace's skills/ directory instead of the shared /app/skills. The same skill may be listed in multiple workspaces (paths are scoped); duplicates within one workspace's list are rejected. Changing a workspace's skills triggers a pod rollout, and the rollout hash is keyed by workspace name, so moving a skill between workspaces rolls out too.
Plugin installation
Install plugins declaratively. The operator runs a dedicated init container that installs each plugin into ~/.openclaw/extensions/ before the agent starts, where is the unscoped npm package basename (so @openclaw/brave-plugin becomes ~/.openclaw/extensions/brave-plugin/):
spec:
plugins:
- "@martian-engineering/lossless-claw"
- "some-other-plugin"
Plugin entries are resolved through the OpenClaw CLI's ClawHub installer, not raw npm install. An optional npm: prefix is accepted for compatibility and stripped before installation, so npm:@scope/plugin and @scope/plugin both run as openclaw plugins install clawhub:@scope/plugin. Use spec.skills when you need npm package source selection for skills.
This is the layout the OpenClaw gateway's plugin discovery expects - it scans direct subdirectories of ~/.openclaw/extensions/ for plugin manifests and skips node_modules/ entirely. The init container shells out to openclaw plugins install clawhub: so plugins published with workspace:* dependency markers, such as the first-party @openclaw/matrix, resolve correctly. Raw npm install rejects those with EUNSUPPORTEDPROTOCOL.
npm lifecycle scripts are disabled globally on the init container (NPM_CONFIG_IGNORE_SCRIPTS=true) to mitigate supply chain attacks. The PVC backs ~/.openclaw/, so installs persist across pod restarts.
If you previously worked around the install-path bug by addingplugins.load.pathsentries to your gateway config (pointing at~/.openclaw/node_modules/), that workaround is no longer needed and can be removed - plugins now land in the documented location and are auto-discovered.
Workspace seeding
Pre-populate the agent workspace with files and directories before the agent starts. Files can be provided inline or referenced from an external ConfigMap -- ideal for GitOps workflows where workspace content is managed alongside your manifests.
Inline files:
spec:
workspace:
initialDirectories:
- tools/scripts
initialFiles:
README.md: |
# My Workspace
This workspace is managed by OpenClaw.
agents/AGENT.md: | # nested paths are supported
# Agent
skills/redmine/SKILL.md: |
# Skill
Keys may contain / for nested files; the operator encodes them for ConfigMap storage and recreates the directory layout when seeding the workspace. The same path safety rules as initialDirectories apply (no leading /, no .., no segment starting with .).
External ConfigMap reference:
spec:
workspace:
configMapRef:
name: my-workspace-files # all keys become workspace files
initialFiles: # inline files (override configMapRef)
EXTRA.md: "additional content"
All keys in the referenced ConfigMap are written as files into the workspace directory. When both configMapRef and initialFiles are specified, inline files take precedence over ConfigMap entries with the same filename.
Merge priority (highest wins): operator-injected files > inline initialFiles > external configMapRef > skill packs.
File update policy
Workspace files are seed-once by default: once a destination exists on persistent storage, later source changes never replace it. That is correct for runtime-owned state the agent writes to, and wrong for files a Git source should keep converging (AGENTS.md, BOUNDARIES.md, runbooks, policy files).
fileUpdatePolicy makes that choice explicit, per workspace or per file:
spec:
workspace:
fileUpdatePolicy: CreateOnly # default for this workspace
configMapRef:
name: main-workspace
managedFiles:
- path: AGENTS.md # updatePolicy defaults to Replace
- path: docs/BOUNDARIES.md
updatePolicy: Replace
- path: STATE.md # pin one file back to seed-once
updatePolicy: CreateOnly
additionalWorkspaces:
- name: print
configMapRef:
name: print-workspace
fileUpdatePolicy: Replace # inherits the top-level default when unset
CreateOnly (the default) keeps the existing behavior. Replace makes a file converge to its source.
What Replace does with local edits. The operator records the hash of the content it last applied, in a marker under /data/.workspace-managed/. A file is rewritten only when that hash moves — that is, when the source genuinely changes. An edit made in the running workspace therefore survives until the next real source change, rather than being wiped on every restart. status.managedResources.workspaceFiles reports the resolved policy and current source hash per path, so kubectl get openclawinstance shows why a file was or was not rewritten.
Guarantees for Replace:
- only explicitly declared files are replaced; a workspace directory is never recursively replaced or pruned
- a destination is never deleted because a source key was removed
- symlink and non-regular destinations are refused, never followed
- writes use a temp file plus atomic rename, with deterministic permissions (
0644) - absolute paths and
..traversal are rejected by the CRD schema and the validating webhook - a path listed in
managedFilesis managed even if no source provides it yet, so adding it to aconfigMapReflater takes effect without a CR change
managedFiles without an updatePolicy means Replace — listing it is an explicit ownership statement. An additionalWorkspaces[].fileUpdatePolicy that is unset inherits the top-level default rather than defaulting independently, so the two cannot drift apart.
Operator-injected files (ENVIRONMENT.md, BOOTSTRAP.md, self-configure files) and skill-pack files are unaffected by this setting -- they have their own lifecycles (bootstrap.enabled and skillPackUpdatePolicy).
Disable operator-managed BOOTSTRAP.md:
BOOTSTRAP.md is seeded on first boot to guide first-run agent onboarding (identity, user preferences, persona). OpenClaw deletes the file after applying it, so on every pod restart or config change the init container would re-copy it and the agent would re-run bootstrap. Opt out once bootstrap is done:
spec:
workspace:
bootstrap:
enabled: false
Defaults to true. ENVIRONMENT.md, self-configure files, and skill-pack files are not affected.
The operator sets a WorkspaceReady status condition to False when the referenced ConfigMap is missing or contains invalid filenames, and True once workspace files are seeded successfully. The controller watches external ConfigMaps for changes and re-reconciles automatically.
How it works: Workspace files are seeded once via an init container. The init container copies files from a read-only ConfigMap volume to the PVC. The main container only sees the PVC (writable), so agents can modify their workspace files and changes persist across pod restarts. ConfigMaps are never mounted directly on the main container.
GitOps example with Kustomize:
# kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: my-namespace # must match the instance namespace
generatorOptions:
disableNameSuffixHash: true # required - operator looks up by exact name
configMapGenerator:
- name: my-workspace-files
files:
- workspace/SOUL.md
- workspace/AGENT.md
Important: Two kustomize settings are required when usingconfigMapGeneratorwithconfigMapRef:
-disableNameSuffixHash: true-- The operator looks up ConfigMaps by exact name. Kustomize's default hash suffix (e.g.-57k7g4dthc) would cause aConfigMapNotFounderror.
-namespace-- Generated ConfigMaps must be in the same namespace as the instance. Without this, kustomize creates them in thedefaultnamespace.
Additional workspaces (multi-agent):
When running multiple agents with isolated workspaces, use additionalWorkspaces to seed files for each agent. Each entry seeds to ~/.openclaw/workspace- -- set matching paths in spec.config.raw.agents.list[].workspace.
spec:
workspace:
configMapRef:
name: main-agent-workspace
additionalWorkspaces:
- name: scheduler
configMapRef:
name: scheduler-workspace
initialFiles:
SOUL.md: "I am the scheduler agent"
initialDirectories:
- tools
config:
raw:
agents:
list:
- id: main
name: "Main Agent"
- id: scheduler
name: "Scheduler Agent"
bindings:
- agentId: scheduler
match:
channel: discord
peer:
kind: channel
id: "123456789" # bind to a specific channel
Each additional workspace supports the same configMapRef, initialFiles, initialDirectories, fileUpdatePolicy, and managedFiles as the default workspace, plus a skills list for workspace-scoped skill installation (see Skill packs). Operator-injected ENVIRONMENT.md is included; BOOTSTRAP.md is not (only the default agent runs onboarding). Max 10 additional workspaces.
Seed-once behavior: Workspace files (both default and additional) are only written on first boot when they don't already exist on the PVC. If an agent modifies its own SOUL.md or AGENT.md at runtime, those changes persist across pod restarts and are never overwritten by the ConfigMap content. To re-seed a file, delete it from the PVC first.
Full GitOps example with multiple agents:
# kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: my-namespace
generatorOptions:
disableNameSuffixHash: true
resources:
- instance.yaml
configMapGenerator:
- name: main-agent-workspace
files:
- agents/main/SOUL.md
- agents/main/AGENT.md
- name: scheduler-workspace
files:
- agents/scheduler/SOUL.md
- agents/scheduler/TOOLS.md
Self-configure
Allow agents to modify their own configuration by creating OpenClawSelfConfig resources via the K8s API. The operator validates each request against the instance's allowedActions policy before applying changes:
spec:
selfConfigure:
enabled: true
allowedActions:
- skills # add/remove skills
- config # patch openclaw.json
- workspaceFiles # add/remove workspace files
- envVars # add/remove environment variables
When enabled, the operator:
- Grants the instance's ServiceAccount RBAC permissions to read its own CRD and create
OpenClawSelfConfigresources - Enables SA token automounting so the agent can authenticate with the K8s API
- Injects a
SELFCONFIG.mdskill file andselfconfig.shhelper script into the workspace - Opens port 6443 egress in the NetworkPolicy for K8s API access
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawSelfConfig
metadata:
name: add-fetch-skill
spec:
instanceRef: my-agent
addSkills:
- "@anthropic/mcp-server-fetch"
The operator validates the request, applies it to the parent OpenClawInstance, and sets the request's status to Applied, Denied, or Failed. Terminal requests are auto-deleted after 1 hour.
GitOps Coexistence
SelfConfig uses Kubernetes Server-Side Apply (SSA) with the field manager name openclaw-selfconfig. This enables safe coexistence with GitOps controllers (FluxCD, ArgoCD, etc.) that manage the same OpenClawInstance resource:
- Per-item ownership -- Skills (set items), env vars (map items by name), and workspace files (map fields) are tracked individually. A SelfConfig can add or remove only the items it owns without conflicting with items managed by other controllers.
- Atomic ownership -- The
config.rawfield is owned atomically. If a GitOps controller also managesconfig.raw,ForceOwnershiptransfers ownership to the SelfConfig field manager on apply. - Removal safety -- When a SelfConfig attempts to remove an item owned by another field manager, the operator emits a
Warning/SelfConfigSkippedRemovalevent identifying the owning manager and includes the warning in the status message. - Non-SSA users are unaffected -- If you do not use
selfConfigure, no SSA field managers are created and existing workflows remain unchanged.
OpenClawSelfConfig CRD spec and spec.selfConfigure fields.
Persistent storage
By default the operator creates a 10Gi PVC and retains it when the CR is deleted (orphan behavior). Override size, storage class, or retention:
spec:
storage:
persistence:
size: 20Gi
storageClass: fast-ssd
orphan: true # default -- PVC is RETAINED when the CR is deleted
# orphan: false -- PVC is deleted with the CR (garbage collected)
To reuse an existing PVC (e.g., after restoring from a backup):
spec:
storage:
persistence:
existingClaim: my-agent-data
Retention is stateful data protection. Because agent workspaces contain irreplaceable data such as memory, notebooks, and conversation history, the default isorphan: true. To re-attach a retained PVC to a new instance, setexistingClaimto its name.
Data volume ownership
Kubernetes applies fsGroup to the root of a volume but never changes its owner, so on most PVCs the directory mounted at ~/.openclaw is owned by root with only the group set to the pod's GID. OpenClaw 2026.9 and later tighten directory modes on that path when they write openclaw.json, and chmod on a directory you do not own fails with EPERM: operation not permitted, fchmod even when you are in its group. Without a fix the gateway crash-loops after the upgrade and openclaw doctor --fix cannot complete.
The operator therefore runs an init-data-owner init container first on every pod start. It runs as root with every capability dropped except CHOWN, changes the owner of the volume root to the pod's runAsUser:runAsGroup when it differs, and exits without touching anything when ownership is already correct. Children of the volume root are created by the pod UID and are never modified.
If your cluster forbids root init containers (for example the restricted Pod Security Standard), disable it and fix ownership out of band once per PVC:
spec:
storage:
fixOwnership: false
Cluster-wide defaults (air-gapped / restricted networks)
For deployments where every instance needs the same registry mirror or the same package-mirror env vars (China regions, air-gapped clusters, private registries), set defaults once on a singleton OpenClawClusterDefaults resource and the operator will merge them into every OpenClawInstance at reconcile time:
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawClusterDefaults
metadata:
name: cluster # name MUST be "cluster" - other names are ignored
spec:
registry: "<account>.dkr.ecr.<region>.amazonaws.com.cn"
env:
- name: NPM_CONFIG_REGISTRY
value: https://registry.npmmirror.com
- name: PIP_INDEX_URL
value: https://mirrors.aliyun.com/pypi/simple/
runtimeDeps:
python: true
Precedence: per-instance fields always win. A cluster default is only applied when the corresponding instance field is unset. For spec.env, cluster-default entries appe
... (README truncated for length)