Profile
Back to NewsBack
GitHub Trending 32 min
Reader Mode
stubbi/openclaw-operator: Kubernetes operator for deploying and managing OpenClaw AI agent instances with production-grade security, observability, and lifecycle management.

stubbi/openclaw-operator: Kubernetes operator for deploying and managing OpenClaw AI agent instances with production-grade security, observability, and lifecycle management.

9 hours ago

OpenClaw Kubernetes Operator — OpenClaws sailing the Kubernetes seas

OpenClaw Kubernetes Operator

License</a> Go Report Card</a> CI</a> Kubernetes</a> Go</a>

Self-host OpenClaw AI agents on Kubernetes with production-grade security, observability, and lifecycle management.

OpenClaw is an AI agent platform that acts on your behalf across Telegram, Discord, WhatsApp, and Signal. It manages your inbox, calendar, smart home, and more through 50+ integrations. While Paperclip Inc. offers fully managed hosting, this operator lets you run OpenClaw on your own infrastructure with the same operational rigor.


Why an Operator?

Deploying AI agents to Kubernetes involves more than a Deployment and a Service. You need network isolation, secret management, persistent storage, health monitoring, optional browser automation, and config rollouts, all wired correctly. This operator encodes those concerns into a single OpenClawInstance custom resource so you can go from zero to production in minutes:

apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawInstance
metadata:
  name: my-agent
spec:
  envFrom:
    - secretRef:
        name: openclaw-api-keys
  storage:
    persistence:
      enabled: true
      size: 10Gi

The operator reconciles this into a fully managed stack of 9+ Kubernetes resources: secured, monitored, and self-healing.

Agents That Adapt Themselves

Agents can autonomously install skills, patch their config, add environment variables, and seed workspace files - all through the Kubernetes API, validated by the operator on every request.

# 1. Enable self-configure on the instance
spec:
  selfConfigure:
    enabled: true
    allowedActions: [skills, config, envVars, workspaceFiles]
# 2. The agent creates this to install a skill at runtime
apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawSelfConfig
metadata:
  name: add-fetch-skill
spec:
  instanceRef: my-agent
  addSkills:
    - "@anthropic/mcp-server-fetch"

Every request is validated against the instance's allowlist policy. Protected config keys cannot be overwritten, and denied requests are logged with a reason. See Self-configure for details.

Note: Without selfConfigure enabled, config or skill changes made by the agent inside the container won't trigger a pod restart. You'll need to restart the pod manually (e.g. kubectl delete pod ) for changes to take effect.

Features

| | Feature | Details | |---|---|---| | Declarative | Single CRD | One resource defines the entire stack: StatefulSet, Service, RBAC, NetworkPolicy, PVC, PDB, Ingress, and more | | Adaptive | Agent self-configure | Agents autonomously install skills, patch config, and adapt their environment via the K8s API - every change validated against an allowlist policy | | Secure | Hardened by default | Non-root (UID 1000), read-only root filesystem, all capabilities dropped, seccomp RuntimeDefault, default-deny NetworkPolicy, validating webhook | | Observable | Built-in metrics | Prometheus metrics, ServiceMonitor integration, structured JSON logging, Kubernetes events | | Flexible | Provider-agnostic config | Use any AI provider (Anthropic, OpenAI, or others) via environment variables and inline or external config | | Config Modes | Merge or overwrite | overwrite replaces config on restart; merge deep-merges with PVC config, preserving runtime changes. Config is restored on every container restart via init container. | | Force Paths | Operator-owned paths under merge | config.forcePaths lists dot-paths the init container rebuilds from the CR on every restart even under mergeMode: merge -- lets managed deployers keep operator-owned config (auth, allowed providers, sandbox image) immune to tenant edits while user-owned config persists | | Skills | Declarative install | Install ClawHub skills, npm packages, or GitHub-hosted skill packs via spec.skills - supports npm: and pack: prefixes. Additional workspaces can declare their own workspace-scoped skills via additionalWorkspaces[].skills | | Plugins | Declarative install | Install OpenClaw plugins via spec.plugins - resolved through the OpenClaw CLI ClawHub installer in a secure init container | | Runtime Deps | pnpm & Python/uv | Built-in init containers install pnpm (via corepack) or Python 3.12 + uv for MCP servers and skills | | Auto-Update | OCI registry polling | Opt-in version tracking: checks the registry for new semver releases, backs up first, rolls out, and auto-rolls back if the new version fails health checks | | Scalable | Auto-scaling | HPA integration with CPU and memory metrics, min/max replica bounds, automatic StatefulSet replica management | | Operational | Instance suspension | Scale to zero with spec.suspended: true - all non-runtime resources remain managed, resume instantly with false | | Resilient | Self-healing lifecycle | PodDisruptionBudgets, health probes, automatic config rollouts via content hashing, 5-minute drift detection | | Disk-Aware Readiness | Opt-in ENOSPC guard | spec.probes.diskReadiness renders the readiness probe as an exec check that ANDs the gateway /readyz signal with a workspace writability + free-space check, so a full or read-only PVC drains the pod from Service endpoints instead of silently accepting writes it cannot persist. Liveness/startup stay HTTP so a full disk never turns into a CrashLoopBackOff. Defaulted off. | | Backup/Restore | S3-backed snapshots | Automatic backup to S3-compatible storage on deletion, pre-update, and on a cron schedule; restore into a new instance from any snapshot | | Workspace Seeding | Initial files & dirs | Pre-populate the workspace with files and directories before the agent starts; reference an external ConfigMap for GitOps workflows | | Gateway Auth | Auto-generated tokens | Automatic shared-secret gateway authentication with a persistent token Secret per instance | | Tailscale | Tailnet access | Expose via Tailscale Serve or Funnel with SSO auth - no Ingress needed | | Extensible | Sidecars & init containers | Chromium for browser automation, Ollama for local LLMs, Tailscale for tailnet access, plus custom init containers and sidecars | | Cloud Native | SA annotations & CA bundles | AWS IRSA / GCP Workload Identity via ServiceAccount annotations; CA bundle injection for corporate proxies | | Cluster Defaults | Singleton CR | OpenClawClusterDefaults (name cluster) fills in unset instance fields - ideal for air-gapped / China regions where every instance would otherwise duplicate the same registry + mirror env boilerplate. Per-instance fields always win. | | Zombie Reaping | Shared PID namespace | spec.shareProcessNamespace defaults to true so the pause container becomes PID 1 and reaps defunct helper processes from QMD, git, plugins, and shells - no custom init image needed |

Architecture

+-----------------------------------------------------------------+
|  OpenClawInstance CR          OpenClawSelfConfig CR              |
|  (your declarative config)   (agent self-modification requests) |
+---------------+-------------------------------------------------+
                | watch
                v
+-----------------------------------------------------------------+
|  OpenClaw Operator                                              |
|  +-----------+  +-------------+  +----------------------------+ |
|  | Reconciler|  |   Webhooks  |  |   Prometheus Metrics       | |
|  |           |  |  (validate  |  |  (reconcile count,         | |
|  |  creates ->  |   & default)|  |   duration, phases)        | |
|  +-----------+  +-------------+  +----------------------------+ |
+---------------+-------------------------------------------------+
                | manages
                v
+-----------------------------------------------------------------+
|  Managed Resources (per instance)                               |
|                                                                 |
|  ServiceAccount -> Role -> RoleBinding    NetworkPolicy         |
|  ConfigMap        PVC      PDB            ServiceMonitor        |
|  GatewayToken Secret                                            |
|                                                                 |
|  StatefulSet                                                    |
|  +-----------------------------------------------------------+ |
|  | Init: config -> pnpm -> python -> skills* -> custom      | |
|  |                                        (* = opt-in)        | |
|  +------------------------------------------------------------+ |
|  | OpenClaw Container  Gateway Proxy (nginx)                  | |
|  |                     Chromium (opt) / Ollama (opt)          | |
|  |                     Tailscale (opt) + custom sidecars      | |
|  +------------------------------------------------------------+ |
|                                                                 |
|  Service (default: 18789, 18793 or custom) -> Ingress (opt)     |
+-----------------------------------------------------------------+

Quick Start

Prerequisites

  • Kubernetes 1.28+
  • Helm 3

1. Install the operator

helm install openclaw-operator \
  oci://ghcr.io/paperclipinc/charts/openclaw-operator \
  --namespace openclaw-operator-system \
  --create-namespace

Alternative: install with Kustomize

# Install CRDs
make install

Deploy the operator

make deploy IMG=ghcr.io/paperclipinc/openclaw-operator:latest

Restrict the operator to specific namespaces

To run the operator with namespaced RBAC instead of cluster-wide permissions, list the namespaces it should watch. The chart switches the namespace-scoped permissions from a ClusterRole/ClusterRoleBinding to per-namespace Role/RoleBinding, and passes --watch-namespaces to the operator so its informer cache is scoped to that list. The operator's own namespace is added to the Secret informer only, so it can still read its backup credentials, and the chart renders a matching Secret-only Role there; no other resource type is watched or granted in the operator namespace. A ClusterRole/ClusterRoleBinding is still created for the cluster-scoped OpenClawClusterDefaults resource, which the operator watches regardless of namespace scoping -- a namespaced Role cannot grant access to a cluster-scoped resource.

helm install openclaw-operator \
  oci://ghcr.io/paperclipinc/charts/openclaw-operator \
  --namespace openclaw-operator-system \
  --create-namespace \
  --set 'watchNamespaces={team-a,team-b}'

Each listed namespace must already exist; the chart does not create them.

To bring your own RBAC entirely (e.g. managed by a separate controller or SecurityCenter policy), disable chart-managed RBAC:

helm install openclaw-operator \
  oci://ghcr.io/paperclipinc/charts/openclaw-operator \
  --namespace openclaw-operator-system \
  --create-namespace \
  --set rbac.create=false

The kubebuilder markers in internal/controller/ and the manager rules helper at charts/openclaw-operator/templates/_helpers.tpl document the minimum permission set the operator requires.

2. Create a secret with your API keys

apiVersion: v1
kind: Secret
metadata:
  name: openclaw-api-keys
type: Opaque
stringData:
  ANTHROPIC_API_KEY: "sk-ant-..."

3. Deploy an OpenClaw instance

apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawInstance
metadata:
  name: my-agent
spec:
  envFrom:
    - secretRef:
        name: openclaw-api-keys
  storage:
    persistence:
      enabled: true
      size: 10Gi
kubectl apply -f secret.yaml -f openclawinstance.yaml

4. Verify

kubectl get openclawinstances

NAME PHASE AGE

my-agent Running 2m

kubectl get pods

NAME READY STATUS AGE

my-agent-0 1/1 Running 2m

Configuration

Inline config (openclaw.json)

spec:
  config:
    raw:
      agents:
        defaults:
          model:
            primary: "anthropic/claude-sonnet-4-20250514"
          sandbox: true
      session:
        scope: "per-sender"

External ConfigMap reference

spec:
  config:
    configMapRef:
      name: my-openclaw-config
      key: openclaw.json

Config changes are detected via SHA-256 hashing and automatically trigger a rolling update. No manual restart needed.

Gateway proxy

By default, each pod includes an nginx reverse proxy sidecar that forwards traffic to the OpenClaw gateway on loopback. Set spec.gateway.enabled: false to disable it:

spec:
  gateway:
    image:
      repository: docker.io/library/nginx
      tag: 1.27-alpine
      # digest: sha256:...  # takes precedence over tag
    resources:
      requests:
        cpu: 10m
        memory: 16Mi
      limits:
        cpu: 100m
        memory: 64Mi

The image and resource fields are optional. Omit them to retain the defaults above, or set an image digest to make the proxy supply-chain reference immutable.

  • Health probes and Service ports target the gateway directly on port 18789
  • gateway.bind is set to 0.0.0.0 instead of loopback
  • The gateway-proxy container and its tmp volume are omitted from the pod
  • To replace the built-in proxy with your own (e.g., Envoy, a signing proxy), disable it and add your proxy via spec.sidecars
  • Warning: Do not set gateway.bind: loopback in your config JSON when the proxy is disabled - the gateway will only listen on 127.0.0.1 with nothing forwarding external traffic, making the pod unreachable. The operator emits a GatewayBindConflict warning event if this misconfiguration is detected.
  • TLS: When the proxy is disabled, the gateway serves plaintext ws:// on 0.0.0.0. Ensure your replacement proxy or Ingress handles TLS termination to avoid exposing unencrypted WebSocket traffic (CWE-319).

Disk-aware readiness

By default the readiness probe is an HTTP GET /readyz against the gateway. For PVC-backed instances, /readyz can stay green while the workspace volume is full or read-only (ENOSPC), so the pod keeps receiving traffic while workspace writes fail. Enable the opt-in disk-aware readiness guard to turn the readiness probe into an exec check that combines the gateway /readyz signal with a workspace writability and free-space check:

spec:
  probes:
    diskReadiness:
      enabled: true            # default: false (existing deployments are unchanged when unset)
      path: /home/openclaw/.openclaw   # optional; defaults to the workspace data mount
      minFree: 128Mi           # optional; minimum free space, a Kubernetes quantity (default 64Mi)
  • When enabled, the readiness probe becomes sh -c exec script that (a) verifies path is writable (test -w), (b) checks free space with df against minFree, then (c) defers to the gateway /readyz on the same loopback port the HTTP probe would use. The pod is Ready only if all checks pass; the script fails closed (non-zero exit) on any failure.
  • Liveness and startup stay HTTP-only (GET /healthz), so a full PVC yields NotReady (draining the pod from Service endpoints) rather than a restart loop / CrashLoopBackOff.
  • The exec script uses only POSIX sh, test, df, and awk. The /readyz HTTP call uses curl or wget if present and is skipped gracefully if neither is in the image, so a missing HTTP client never makes a healthy pod permanently NotReady (disk checks still run).
  • This is a secondary, defense-in-depth guard; the application-level /readyz endpoint remains the primary readiness signal.

Gateway authentication

The operator automatically generates a gateway token Secret for each instance and injects it into both the config JSON (gateway.auth.mode: token) and the OPENCLAW_GATEWAY_TOKEN env var. The token authenticates the gateway connection. A browser's Control UI device identity and one-time pairing are separate security checks.

  • The token is generated once and never overwritten - rotate it by editing the Secret directly
  • If you set gateway.auth.token in your config or OPENCLAW_GATEWAY_TOKEN in spec.env, your value takes precedence
  • To bring your own token Secret, set spec.gateway.existingSecret - the operator will use it instead of auto-generating one (the Secret must have a key named token)
  • The operator sets OPENCLAW_DISABLE_BONJOUR=1 because mDNS discovery is not useful in Kubernetes. This does not disable Control UI device identity.
  • The operator sets gateway.mode: local, which current OpenClaw releases require for a gateway that owns local state. Exposure remains controlled by the Service, Ingress, mesh, and gateway bind settings.
  • Current OpenClaw releases ignore the retired gateway.controlUi.dangerouslyDisableDeviceAuth setting, so the operator does not emit it.
  • On the first browser connection, approve the pending device once from an administrative workstation:
kubectl exec -n <namespace> <instance>-0 -c openclaw -- openclaw devices list
  kubectl exec -n <namespace> <instance>-0 -c openclaw -- openclaw devices approve <request-id>
  • Supplying the gateway token, including in the Control UI, does not replace browser device pairing.
  • Since v2026.2.24, OpenClaw restricts gateway.allowedOrigins to same-origin by default - if accessing via a non-default hostname (e.g. Ingress), set gateway.allowedOrigins: ["*"] in your config

Control UI allowed origins

The operator auto-injects gateway.controlUi.allowedOrigins so the Control UI works through reverse proxies without CORS errors. Origins are derived from:

  • Localhost (always): http://localhost:18789, http://127.0.0.1:18789 for port-forwarding
  • Ingress hosts: scheme determined from TLS config (https:// if TLS, http:// otherwise)
  • Explicit extras: spec.gateway.controlUiOrigins for custom proxy URLs
If you set gateway.controlUi.allowedOrigins directly in your config JSON, the operator will not override it.

Chromium sidecar

Enable headless browser automation for web scraping, screenshots, and browser-based integrations:

spec:
  chromium:
    enabled: true
    image:
      repository: chromedp/headless-shell  # default
      tag: "stable"
    resources:
      requests:
        cpu: "250m"
        memory: "512Mi"
      limits:
        cpu: "1000m"
        memory: "2Gi"
    # Pass extra flags to the Chromium process (appended to built-in anti-bot defaults)
    extraArgs:
      - "--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
    # Inject extra environment variables into the sidecar
    extraEnv:
      - name: DISPLAY
        value: ":99"

When enabled, the operator automatically:

  • Injects a CHROMIUM_URL environment variable into the main container
  • Configures browser profiles in the OpenClaw config - both "default" and "chrome" profiles are set to point at the sidecar's CDP endpoint, so browser tool calls work regardless of which profile name the LLM passes
  • Sets up shared memory, security contexts, and health probes for the sidecar
  • Applies anti-bot-detection flags by default (--disable-blink-features=AutomationControlled, --disable-features=AutomationControlled, --no-first-run)

Persistent browser profiles

By default, all browser state (cookies, localStorage, session tokens) is lost on pod restart. Enable persistence to retain browser profiles across restarts:

spec:
  chromium:
    enabled: true
    persistence:
      enabled: true          # default: false
      storageClass: ""        # optional - uses cluster default if empty
      size: "1Gi"             # default: 1Gi
      existingClaim: ""       # optional - use a pre-existing PVC

When persistence is enabled, the operator creates a dedicated PVC and passes --user-data-dir=/chromium-data to Chrome so that cookies, localStorage, IndexedDB, cached credentials, and session tokens survive pod restarts. This is useful for authenticated browser automation, MFA-protected services, and long-running browser workflows.

Security note: Persistent browser profiles contain sensitive session tokens. The PVC has the same security posture as other instance volumes. Ensure your StorageClass supports encryption at rest for sensitive workloads.

Ollama sidecar

Run local LLMs alongside your agent for private, low-latency inference without external API calls:

spec:
  ollama:
    enabled: true
    models:
      - llama3.2
      - nomic-embed-text
    gpu: 1
    storage:
      sizeLimit: 30Gi
    resources:
      requests:
        cpu: "1"
        memory: "4Gi"
      limits:
        cpu: "4"
        memory: "16Gi"

When enabled, the operator:

  • Injects an OLLAMA_HOST environment variable into the main container
  • Pre-pulls specified models via an init container before the agent starts
  • Configures GPU resource limits when gpu is set (nvidia.com/gpu)
  • Mounts a model cache volume (emptyDir by default, or an existing PVC via storage.existingClaim)
See Custom AI Providers for configuring OpenClaw to use Ollama models via environment variables.

Web terminal sidecar

Provide browser-based shell access to running instances for debugging and inspection without requiring kubectl exec:

spec:
  webTerminal:
    enabled: true
    readOnly: false
    credential:
      secretRef:
        name: my-terminal-creds
    resources:
      requests:
        cpu: "50m"
        memory: "64Mi"
      limits:
        cpu: "200m"
        memory: "128Mi"

When enabled, the operator:

  • Injects a ttyd sidecar container on port 7681
  • Mounts the instance data volume at /home/openclaw/.openclaw so you can inspect config, logs, and data files
  • Adds the web terminal port to the Service and NetworkPolicy for external access
  • Supports basic auth via a Secret with username and password keys
  • Supports read-only mode (readOnly: true) for production environments where shell input should be disabled

Tailscale integration

Expose your instance via Tailscale Serve (tailnet-only) or Funnel (public internet) - no Ingress or LoadBalancer needed:

spec:
  tailscale:
    enabled: true
    mode: serve          # "serve" (tailnet only) or "funnel" (public internet)
    authKeySecretRef:
      name: tailscale-auth
    authSSO: true        # allow passwordless login for tailnet members
    hostname: my-agent   # defaults to instance name
    image:
      repository: ghcr.io/tailscale/tailscale  # default
      tag: latest
    resources:
      requests:
        cpu: 50m
        memory: 64Mi
      limits:
        cpu: 200m
        memory: 256Mi

When enabled, the operator runs a Tailscale sidecar (tailscaled) that handles serve/funnel declaratively via TS_SERVE_CONFIG. An init container copies the tailscale CLI binary to a shared volume so the main container can call tailscale whois for SSO authentication. The sidecar runs in userspace mode (TS_USERSPACE=true) - no NET_ADMIN capability needed.

State persistence: Tailscale node identity and TLS certificates are automatically persisted to a Kubernetes Secret (-ts-state) via TS_KUBE_SECRET. This prevents hostname incrementing (device-1, device-2, ...) and Let's Encrypt certificate re-issuance across pod restarts. The operator pre-creates the state Secret, grants the pod's ServiceAccount get/update/patch access to it, and mounts the SA token automatically.

Use ephemeral+reusable auth keys from the Tailscale admin console. When authSSO is enabled, tailnet members can authenticate without a gateway token.

NetBird integration

NetBird is a self-hostable alternative to Tailscale: the same WireGuard data plane, with a control plane you can run yourself.

spec:
  netbird:
    enabled: true
    setupKeySecretRef:
      name: netbird-setup-key       # key: "setupkey"
    managementURL: https://netbird.example.com:33073   # omit for NetBird's hosted control plane
    hostname: my-agent              # defaults to the instance name

Use a reusable, ephemeral setup key from the NetBird dashboard. The operator injects it as NB_SETUP_KEY from the referenced Secret -- it is never written into the pod spec as a literal -- and rolls the pod when the Secret changes.

The sidecar runs in netstack (userspace) mode, so it keeps the same Restricted PSS posture as every other container the operator builds: all capabilities dropped, read-only root filesystem, non-root, seccomp RuntimeDefault. A kernel-mode peer would need NET_ADMIN and /dev/net/tun, which is a different security decision than this operator makes by default.

Peer state lives on an emptyDir, so the peer re-enrolls on restart (which a reusable setup key handles). Unlike Tailscale, NetBird needs no Kubernetes API access, so no ServiceAccount token is mounted and no state Secret is created.

Mesh providers are mutually exclusive. Enabling both tailscale and netbird is rejected by the validating webhook: two overlay clients in one pod would race for the same egress rules and the agent's routing.

Feature comparison:

| | Tailscale | NetBird | |---|---|---| | Self-hostable control plane | no | yes (managementURL) | | Credential | auth key | setup key | | Serve/Funnel ingress | yes (mode) | not applicable | | Gateway SSO (authSSO) | yes | no identity header equivalent | | Needs Kubernetes API | yes (state Secret) | no | | Persistent node identity | yes (state Secret) | re-enrolls on restart |

Both are implementations of one internal MeshProvider interface, so a third provider means implementing that interface and adding one table entry -- not another copy of the StatefulSet, NetworkPolicy, RBAC and config-enrichment paths.

Config merge mode

By default, the operator overwrites the config file on every pod restart. Set mergeMode: merge to deep-merge operator config with existing PVC config, preserving runtime changes made by the agent:

spec:
  config:
    mergeMode: merge
    raw:
      agents:
        defaults:
          model:
            primary: "anthropic/claude-sonnet-4-20250514"

Caveat: In merge mode, removing a key from the CR does not remove it from the PVC config - the old value persists because deep-merge only adds or updates keys. If you need to remove a stale config key, temporarily switch to mergeMode: overwrite, apply, wait for the pod to restart, then switch back to merge.

Partial overwrite under merge mode (forcePaths)

Under mergeMode: merge the operator preserves runtime changes the agent (or a tenant via the Control UI) wrote into the config file. For managed multi-tenant deployments this is a problem: a tenant can persist arbitrary values into operator-owned subtrees -- for example models.providers..apiKey -- and route inference through their own third-party key while consuming the deployer's compute.

spec.config.forcePaths is the partial-overwrite escape hatch. For each listed dot-path the init container deletes that subtree from the PVC config and re-applies it from spec.config.raw on every pod restart, so listed paths always match the CR while everything else still persists.

spec:
  config:
    mergeMode: merge
    forcePaths:
      - gateway
      - models.providers
      - agents.defaults.sandbox
    raw:
      gateway:
        auth:
          mode: token
      models:
        providers:
          openai:
            baseUrl: "https://api.openai.com"
      agents:
        defaults:
          sandbox: true

With the above, channels., settings., and any other user-owned path the agent writes via the Control UI persists across pod restarts; gateway., models.providers., and agents.defaults.sandbox are rebuilt from the CR on every reconcile.

forcePaths is only valid under mergeMode: merge. The validating webhook rejects it under overwrite (where the whole file is already rebuilt every restart) and rejects malformed paths (empty segments, leading or trailing dot, characters outside [a-zA-Z0-9._-]). The same logic runs in both the init container (on pod restart) and the postStart lifecycle hook (on container restart without pod recreation), so an attacker cannot bypass the contract by triggering one form of restart over another.

Skill installation

Install skills declaratively. The operator runs an init container that fetches each skill before the agent starts. Entries use ClawHub by default, or prefix with npm: to install from npmjs.com. ClawHub installs are idempotent - if a skill is already installed (e.g., when using persistent storage), it is skipped rather than failing:

spec:
  skills:
    - "@anthropic/mcp-server-fetch"       # ClawHub (default)
    - "npm:@openclaw/matrix"              # npm package from npmjs.com

npm lifecycle scripts are disabled globally on the init container (NPM_CONFIG_IGNORE_SCRIPTS=true) to mitigate supply chain attacks.

Skill packs

Skill packs bundle multiple files (SKILL.md, scripts, config) into a single installable unit hosted on GitHub. Use the pack: prefix with owner/repo/path format:

spec:
  skills:
    - "pack:paperclipinc/skills/image-gen"            # latest from default branch
    - "pack:paperclipinc/skills/[email protected]"     # pinned to tag
    - "pack:myorg/private-skills/custom-tool@main"       # private repo (requires GITHUB_TOKEN)

Packs are resolved in one of two modes:

1. Manifest mode (explicit) -- the pack path contains a skillpack.json describing which files to seed and where:

{
  "files": {
    "skills/image-gen/SKILL.md": "SKILL.md",
    "skills/image-gen/scripts/generate.py": "scripts/generate.py"
  },
  "directories": ["skills/image-gen/scripts"],
  "config": {
    "image-gen": {"enabled": true}
  }
}

2. Raw-repo mode (autodiscovery) -- when no skillpack.json is present and the pack path contains a SKILL.md, the operator installs the entire directory verbatim into skills// in the workspace. This is useful for multi-skill repositories like fluxcd/agent-skills that follow a conventional skills//SKILL.md layout without per-skill manifests:

spec:
  skills:
    - "pack:fluxcd/agent-skills/skills/gitops-repo-audit@main"
    # installs every file under skills/gitops-repo-audit/ into the workspace
    # at skills/gitops-repo-audit/ (including nested assets, schemas, etc.)

Raw mode does not inject config entries into config.raw.skills.entries -- use manifest mode if you need that. The operator refuses to install if GitHub truncates the tree response for very large repositories (add a skillpack.json manifest in that case).

The operator resolves packs via the GitHub Contents + Git Trees APIs (cached for 5 minutes), seeds files into the workspace via the init container, and (in manifest mode) injects config entries into config.raw.skills.entries with user overrides taking precedence. Set GITHUB_TOKEN on the operator deployment for private repo access.

Updating pack contents. By default (spec.skillPackUpdatePolicy: Replace), pack-seeded files converge to the declared pack revision on every pod start: changing a pinned @tag/@commit (or pushing to a tracked branch) overwrites the seeded files, and files that are no longer part of any declared pack are removed. The operator tracks what it seeded in a manifest at /data/.skillpack-manifest on the data volume, so user-created workspace files are never touched. Files at pack-declared paths are operator-managed -- local edits to them are reverted on restart. Set spec.skillPackUpdatePolicy: CreateOnly to opt out and keep the legacy seed-once behavior (files are never overwritten or removed after first seeding; updating a pinned revision then has no effect on already-seeded files).

Workspace-scoped skills (multi-agent). Additional workspaces can declare their own skills with the same reference formats as spec.skills:

spec:
  workspace:
    additionalWorkspaces:
      - name: secondary
        skills:
          - "pack:example-org/openclaw-skills/skills/[email protected]"
          - "@acme/browser-use"          # ClawHub, installed into workspace-secondary/skills/
          - "npm:@acme/cli-tool"         # npm binaries are global (~/.local/bin), shared by all agents

pack: entries resolve exactly like top-level packs (private repos via GITHUB_TOKEN, pinned tags/commits) but seed into ~/.openclaw/workspace-/ and are tracked in a per-workspace manifest (/data/.skillpack-manifest-ws-), so skillPackUpdatePolicy applies per workspace. ClawHub entries are installed with clawhub --workdir so the skill lands in that workspace's skills/ directory instead of the shared /app/skills. The same skill may be listed in multiple workspaces (paths are scoped); duplicates within one workspace's list are rejected. Changing a workspace's skills triggers a pod rollout, and the rollout hash is keyed by workspace name, so moving a skill between workspaces rolls out too.

Plugin installation

Install plugins declaratively. The operator runs a dedicated init container that installs each plugin into ~/.openclaw/extensions// before the agent starts, where is the unscoped npm package basename (so @openclaw/brave-plugin becomes ~/.openclaw/extensions/brave-plugin/):

spec:
  plugins:
    - "@martian-engineering/lossless-claw"
    - "some-other-plugin"

Plugin entries are resolved through the OpenClaw CLI's ClawHub installer, not raw npm install. An optional npm: prefix is accepted for compatibility and stripped before installation, so npm:@scope/plugin and @scope/plugin both run as openclaw plugins install clawhub:@scope/plugin. Use spec.skills when you need npm package source selection for skills.

This is the layout the OpenClaw gateway's plugin discovery expects - it scans direct subdirectories of ~/.openclaw/extensions/ for plugin manifests and skips node_modules/ entirely. The init container shells out to openclaw plugins install clawhub: so plugins published with workspace:* dependency markers, such as the first-party @openclaw/matrix, resolve correctly. Raw npm install rejects those with EUNSUPPORTEDPROTOCOL.

npm lifecycle scripts are disabled globally on the init container (NPM_CONFIG_IGNORE_SCRIPTS=true) to mitigate supply chain attacks. The PVC backs ~/.openclaw/, so installs persist across pod restarts.

If you previously worked around the install-path bug by adding plugins.load.paths entries to your gateway config (pointing at ~/.openclaw/node_modules/), that workaround is no longer needed and can be removed - plugins now land in the documented location and are auto-discovered.

Workspace seeding

Pre-populate the agent workspace with files and directories before the agent starts. Files can be provided inline or referenced from an external ConfigMap -- ideal for GitOps workflows where workspace content is managed alongside your manifests.

Inline files:

spec:
  workspace:
    initialDirectories:
      - tools/scripts
    initialFiles:
      README.md: |
        # My Workspace
        This workspace is managed by OpenClaw.
      agents/AGENT.md: |              # nested paths are supported
        # Agent
      skills/redmine/SKILL.md: |
        # Skill

Keys may contain / for nested files; the operator encodes them for ConfigMap storage and recreates the directory layout when seeding the workspace. The same path safety rules as initialDirectories apply (no leading /, no .., no segment starting with .).

External ConfigMap reference:

spec:
  workspace:
    configMapRef:
      name: my-workspace-files      # all keys become workspace files
    initialFiles:                    # inline files (override configMapRef)
      EXTRA.md: "additional content"

All keys in the referenced ConfigMap are written as files into the workspace directory. When both configMapRef and initialFiles are specified, inline files take precedence over ConfigMap entries with the same filename.

Merge priority (highest wins): operator-injected files > inline initialFiles > external configMapRef > skill packs.

File update policy

Workspace files are seed-once by default: once a destination exists on persistent storage, later source changes never replace it. That is correct for runtime-owned state the agent writes to, and wrong for files a Git source should keep converging (AGENTS.md, BOUNDARIES.md, runbooks, policy files).

fileUpdatePolicy makes that choice explicit, per workspace or per file:

spec:
  workspace:
    fileUpdatePolicy: CreateOnly     # default for this workspace
    configMapRef:
      name: main-workspace
    managedFiles:
      - path: AGENTS.md              # updatePolicy defaults to Replace
      - path: docs/BOUNDARIES.md
        updatePolicy: Replace
      - path: STATE.md               # pin one file back to seed-once
        updatePolicy: CreateOnly
    additionalWorkspaces:
      - name: print
        configMapRef:
          name: print-workspace
        fileUpdatePolicy: Replace    # inherits the top-level default when unset

CreateOnly (the default) keeps the existing behavior. Replace makes a file converge to its source.

What Replace does with local edits. The operator records the hash of the content it last applied, in a marker under /data/.workspace-managed/. A file is rewritten only when that hash moves — that is, when the source genuinely changes. An edit made in the running workspace therefore survives until the next real source change, rather than being wiped on every restart. status.managedResources.workspaceFiles reports the resolved policy and current source hash per path, so kubectl get openclawinstance -o yaml shows why a file was or was not rewritten.

Guarantees for Replace:

  • only explicitly declared files are replaced; a workspace directory is never recursively replaced or pruned
  • a destination is never deleted because a source key was removed
  • symlink and non-regular destinations are refused, never followed
  • writes use a temp file plus atomic rename, with deterministic permissions (0644)
  • absolute paths and .. traversal are rejected by the CRD schema and the validating webhook
  • a path listed in managedFiles is managed even if no source provides it yet, so adding it to a configMapRef later takes effect without a CR change
Listing a path in managedFiles without an updatePolicy means Replace — listing it is an explicit ownership statement. An additionalWorkspaces[].fileUpdatePolicy that is unset inherits the top-level default rather than defaulting independently, so the two cannot drift apart.

Operator-injected files (ENVIRONMENT.md, BOOTSTRAP.md, self-configure files) and skill-pack files are unaffected by this setting -- they have their own lifecycles (bootstrap.enabled and skillPackUpdatePolicy).

Disable operator-managed BOOTSTRAP.md:

BOOTSTRAP.md is seeded on first boot to guide first-run agent onboarding (identity, user preferences, persona). OpenClaw deletes the file after applying it, so on every pod restart or config change the init container would re-copy it and the agent would re-run bootstrap. Opt out once bootstrap is done:

spec:
  workspace:
    bootstrap:
      enabled: false

Defaults to true. ENVIRONMENT.md, self-configure files, and skill-pack files are not affected.

The operator sets a WorkspaceReady status condition to False when the referenced ConfigMap is missing or contains invalid filenames, and True once workspace files are seeded successfully. The controller watches external ConfigMaps for changes and re-reconciles automatically.

How it works: Workspace files are seeded once via an init container. The init container copies files from a read-only ConfigMap volume to the PVC. The main container only sees the PVC (writable), so agents can modify their workspace files and changes persist across pod restarts. ConfigMaps are never mounted directly on the main container.

GitOps example with Kustomize:

# kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization

namespace: my-namespace # must match the instance namespace

generatorOptions: disableNameSuffixHash: true # required - operator looks up by exact name

configMapGenerator: - name: my-workspace-files files: - workspace/SOUL.md - workspace/AGENT.md

Important: Two kustomize settings are required when using configMapGenerator with configMapRef:
- disableNameSuffixHash: true -- The operator looks up ConfigMaps by exact name. Kustomize's default hash suffix (e.g. -57k7g4dthc) would cause a ConfigMapNotFound error.
- namespace -- Generated ConfigMaps must be in the same namespace as the instance. Without this, kustomize creates them in the default namespace.

Additional workspaces (multi-agent):

When running multiple agents with isolated workspaces, use additionalWorkspaces to seed files for each agent. Each entry seeds to ~/.openclaw/workspace-/ -- set matching paths in spec.config.raw.agents.list[].workspace.

spec:
  workspace:
    configMapRef:
      name: main-agent-workspace
    additionalWorkspaces:
      - name: scheduler
        configMapRef:
          name: scheduler-workspace
        initialFiles:
          SOUL.md: "I am the scheduler agent"
        initialDirectories:
          - tools
  config:
    raw:
      agents:
        list:
          - id: main
            name: "Main Agent"
          - id: scheduler
            name: "Scheduler Agent"
      bindings:
        - agentId: scheduler
          match:
            channel: discord
            peer:
              kind: channel
              id: "123456789"        # bind to a specific channel

Each additional workspace supports the same configMapRef, initialFiles, initialDirectories, fileUpdatePolicy, and managedFiles as the default workspace, plus a skills list for workspace-scoped skill installation (see Skill packs). Operator-injected ENVIRONMENT.md is included; BOOTSTRAP.md is not (only the default agent runs onboarding). Max 10 additional workspaces.

Seed-once behavior: Workspace files (both default and additional) are only written on first boot when they don't already exist on the PVC. If an agent modifies its own SOUL.md or AGENT.md at runtime, those changes persist across pod restarts and are never overwritten by the ConfigMap content. To re-seed a file, delete it from the PVC first.

Full GitOps example with multiple agents:

# kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization

namespace: my-namespace

generatorOptions: disableNameSuffixHash: true

resources: - instance.yaml

configMapGenerator: - name: main-agent-workspace files: - agents/main/SOUL.md - agents/main/AGENT.md - name: scheduler-workspace files: - agents/scheduler/SOUL.md - agents/scheduler/TOOLS.md

Self-configure

Allow agents to modify their own configuration by creating OpenClawSelfConfig resources via the K8s API. The operator validates each request against the instance's allowedActions policy before applying changes:

spec:
  selfConfigure:
    enabled: true
    allowedActions:
      - skills        # add/remove skills
      - config        # patch openclaw.json
      - workspaceFiles # add/remove workspace files
      - envVars       # add/remove environment variables

When enabled, the operator:

  • Grants the instance's ServiceAccount RBAC permissions to read its own CRD and create OpenClawSelfConfig resources
  • Enables SA token automounting so the agent can authenticate with the K8s API
  • Injects a SELFCONFIG.md skill file and selfconfig.sh helper script into the workspace
  • Opens port 6443 egress in the NetworkPolicy for K8s API access
The agent creates a request like:

apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawSelfConfig
metadata:
  name: add-fetch-skill
spec:
  instanceRef: my-agent
  addSkills:
    - "@anthropic/mcp-server-fetch"

The operator validates the request, applies it to the parent OpenClawInstance, and sets the request's status to Applied, Denied, or Failed. Terminal requests are auto-deleted after 1 hour.

GitOps Coexistence

SelfConfig uses Kubernetes Server-Side Apply (SSA) with the field manager name openclaw-selfconfig. This enables safe coexistence with GitOps controllers (FluxCD, ArgoCD, etc.) that manage the same OpenClawInstance resource:

  • Per-item ownership -- Skills (set items), env vars (map items by name), and workspace files (map fields) are tracked individually. A SelfConfig can add or remove only the items it owns without conflicting with items managed by other controllers.
  • Atomic ownership -- The config.raw field is owned atomically. If a GitOps controller also manages config.raw, ForceOwnership transfers ownership to the SelfConfig field manager on apply.
  • Removal safety -- When a SelfConfig attempts to remove an item owned by another field manager, the operator emits a Warning / SelfConfigSkippedRemoval event identifying the owning manager and includes the warning in the status message.
  • Non-SSA users are unaffected -- If you do not use selfConfigure, no SSA field managers are created and existing workflows remain unchanged.
See the API reference for the full OpenClawSelfConfig CRD spec and spec.selfConfigure fields.

Persistent storage

By default the operator creates a 10Gi PVC and retains it when the CR is deleted (orphan behavior). Override size, storage class, or retention:

spec:
  storage:
    persistence:
      size: 20Gi
      storageClass: fast-ssd
      orphan: true   # default -- PVC is RETAINED when the CR is deleted
      # orphan: false  -- PVC is deleted with the CR (garbage collected)

To reuse an existing PVC (e.g., after restoring from a backup):

spec:
  storage:
    persistence:
      existingClaim: my-agent-data
Retention is stateful data protection. Because agent workspaces contain irreplaceable data such as memory, notebooks, and conversation history, the default is orphan: true. To re-attach a retained PVC to a new instance, set existingClaim to its name.

Data volume ownership

Kubernetes applies fsGroup to the root of a volume but never changes its owner, so on most PVCs the directory mounted at ~/.openclaw is owned by root with only the group set to the pod's GID. OpenClaw 2026.9 and later tighten directory modes on that path when they write openclaw.json, and chmod on a directory you do not own fails with EPERM: operation not permitted, fchmod even when you are in its group. Without a fix the gateway crash-loops after the upgrade and openclaw doctor --fix cannot complete.

The operator therefore runs an init-data-owner init container first on every pod start. It runs as root with every capability dropped except CHOWN, changes the owner of the volume root to the pod's runAsUser:runAsGroup when it differs, and exits without touching anything when ownership is already correct. Children of the volume root are created by the pod UID and are never modified.

If your cluster forbids root init containers (for example the restricted Pod Security Standard), disable it and fix ownership out of band once per PVC:

spec:
  storage:
    fixOwnership: false

Cluster-wide defaults (air-gapped / restricted networks)

For deployments where every instance needs the same registry mirror or the same package-mirror env vars (China regions, air-gapped clusters, private registries), set defaults once on a singleton OpenClawClusterDefaults resource and the operator will merge them into every OpenClawInstance at reconcile time:

apiVersion: openclaw.rocks/v1alpha1
kind: OpenClawClusterDefaults
metadata:
  name: cluster  # name MUST be "cluster" - other names are ignored
spec:
  registry: "<account>.dkr.ecr.<region>.amazonaws.com.cn"
  env:
    - name: NPM_CONFIG_REGISTRY
      value: https://registry.npmmirror.com
    - name: PIP_INDEX_URL
      value: https://mirrors.aliyun.com/pypi/simple/
  runtimeDeps:
    python: true

Precedence: per-instance fields always win. A cluster default is only applied when the corresponding instance field is unset. For spec.env, cluster-default entries appe

... (README truncated for length)

Chat with me