Profile
Back to NewsBack
GitHub Trending 32 min
Reader Mode
rajsinghtech/garage-operator: A Kubernetes operator for managing Garage - a distributed S3-compatible object storage system designed for self-hosting.

rajsinghtech/garage-operator: A Kubernetes operator for managing Garage - a distributed S3-compatible object storage system designed for self-hosting.

4 hours ago

Garage Kubernetes Operator

Garage Kubernetes Operator

S3-Compatible Object Storage on Kubernetes

Read the operator documentation · Releases · Support

CI Go Report Card Latest Release Ask DeepWiki

A Kubernetes operator for Garage - distributed, self-hosted object storage with multi-cluster federation.

  • Declarative cluster lifecycle — StatefulSet, config, and layout managed via CRDs
  • Unified storage + gateway tiers in one CR (v1beta2) — combine durable storage pods and persistent-identity S3 gateways in a single GarageCluster
  • Node-local pools — bind Garage identities to selected Kubernetes Nodes and HostPath disks, including multi-disk layouts
  • Bucket & key management — create buckets, quotas, and S3 credentials with kubectl
  • Multi-cluster federation — span storage across Kubernetes clusters with automatic node discovery
  • Persistent-identity gateway pods — StatefulSet with a small metadata PVC; gateway pods keep the same Garage node identity across restarts and participate in the cluster layout with capacity: null (matching upstream garage layout assign --gateway)
  • Scale subresource — kubectl scale and autoscaler support for the Auto-managed default storage group (and v1beta1 edge gateways)
  • COSI driver — optional Kubernetes-native object storage provisioning

Custom Resources

| CRD | Description | |-----|-------------| | GarageCluster | Deploys and manages a Garage cluster (storage and/or gateway tiers) | | GarageBucket | Creates buckets with quotas and website hosting | | GarageKey | Provisions S3 access keys with per-bucket permissions | | GarageNode | Fine-grained node layout control (zone, capacity, tags) | | GarageAdminToken | Creates static Admin bootstrap material in a namespace-local Secret | | GarageReferenceGrant | Grants selected cross-namespace access to clusters, buckets, and keys |

Install

Requires Kubernetes 1.25+, or **1.27+ if you use node-local pools**, which depend on Pod scheduling gates for their activation fence.

The Helm chart enables admission and conversion webhooks by default, so install cert-manager first. Disabling webhooks is limited to local development or simple v1beta2-only installs that neither use nodeLocalPools nor rely on admission-protected storage deletion. It removes those safety checks and all v1beta1 conversion support; node-local pools and controller-managed persistent claims do not support that mode. EmptyDir remains fully supported there; explicit existingClaim volumes can be mounted, but their PVC-backed rollout and recovery paths remain fenced until admission is enabled. The webhooks also reserve managed PVC finalizer removal to the operator service account, preventing namespace users with PVC update rights from reopening a same-name claim replacement race before StatefulSet ownership is established.

helm install garage-operator oci://ghcr.io/rajsinghtech/charts/garage-operator \
  --namespace garage-operator-system \
  --create-namespace
helm install garage-operator oci://ghcr.io/rajsinghtech/charts/garage-operator \
  --namespace garage-operator-system \
  --create-namespace \
  --set webhooks.enabled=false

Verifying release artifacts

Released container images and Helm charts are signed with cosign keyless signing (the GitHub Actions OIDC identity — no long-lived keys), and carry SLSA build provenance. The image additionally carries an SPDX SBOM. All three are stored in GHCR as OCI referrers of the artifact digest.

IMAGE=ghcr.io/rajsinghtech/garage-operator:v0.8.0

Signature

cosign verify "$IMAGE" \ --certificate-identity-regexp '^https://github.com/rajsinghtech/garage-operator/\.github/workflows/docker\.yml@refs/' \ --certificate-oidc-issuer https://token.actions.githubusercontent.com

Provenance and SBOM

gh attestation verify "oci://$IMAGE" --repo rajsinghtech/garage-operator cosign download attestation "$IMAGE" --predicate-type https://spdx.dev/Document/v2.3

The Helm chart is signed the same way (--certificate-identity-regexp ending in helm\.yml@refs/), and dist/install.yaml attached to each GitHub release has a provenance attestation verifiable with gh attestation verify install.yaml --repo rajsinghtech/garage-operator.

Under a policy controller, pin by digest and require the signature — e.g. Kyverno verifyImages with keyless.issuer: https://token.actions.githubusercontent.com and the subject regexp above.

Garage Version Compatibility

The Garage version is yours to choose — GarageCluster.spec.image, GarageNode.spec.image, or the chart-wide defaultGarageImage. The chart's appVersion tracks the operator, not Garage.

| Operator | Garage minimum | Garage tested in CI | Notes | |---|---|---|---| | 0.8.x | v2.0.0 | v2.4.1 (default, all suites), v2.0.0 (floor lane); nightly main-v2 canary | Admin API v2; node-local pools require Kubernetes 1.27+ | | 0.7.x | v2.0.0 | v2.4.0, v2.2.0 | Admin API v2; node-local pools require Kubernetes 1.27+ | | 0.6.x | v2.0.0 | v2.3.0, v2.2.0 | Admin API v2 only |

dxflrs/garage:v2.4.1@sha256:9c96caa2612d3411acc5b0e6701fb238dbfba33e533a6d7d3d811a4b12d0d020 is the built-in default when spec.image is unset, so default deployments run the multi-platform image index that the Ginkgo suite and the topology suites (multi-cluster, external gateway, IPv6, single-cluster) are pinned to. Two further lanes back the "v2.x range" claim: a floor lane runs the core bucket and key path on the pinned Garage v2.0.0 index (the oldest supported release; v2.0.0 is the only v2.0.x release), and a nightly canary runs the same path on a build of Garage's main-v2 branch (.github/workflows/garage-canary.yml). One lane enables Garage's native kubernetes_discovery with namespaced RBAC. Any other Garage image, including every other release, is compatible by contract (the Admin API v2) but not individually tested; see the compatibility matrix.

Garage 0.x and 1.x are not supported. The operator drives buckets, keys, layout, and repair exclusively through the /v2/... admin API, which first shipped in Garage v2.0.0. Against an older node every admin call 404s and no cluster will reconcile.

Some fields need a newer Garage than the v2.0.0 floor:

| Field | Requires | Behavior on older Garage | |---|---|---| | GarageBucket.spec.lifecycle | v2.3.0 | Older nodes accept the write and drop the field. The operator reads the rules back and sets LifecycleConfigured=False naming this requirement, rather than reporting a success that never took effect. The bucket itself still reconciles. | | GarageCluster.spec.database.engine: fjall, spec.database.fjallBlockCacheSize | v2.1.0 | Unknown config key, silently ignored; Garage falls back to the default engine | | GarageCluster.spec.blocks.maxConcurrentReads | v2.1.0 | Silently ignored | | GarageCluster.spec.blocks.maxConcurrentWritesPerRequest | v2.2.0 | Silently ignored |

Garage's TOML parser ignores unknown keys, so setting a too-new config field degrades to a no-op rather than a crashloop. The operator only emits these keys when you set the corresponding field.

The Garage version each cluster is actually running is reported back on the CR:

kubectl get garagecluster garage -o jsonpath='{.status.buildInfo.version}'

API Versions

GarageCluster is served under two API versions; all other CRDs are v1beta1.

| Version | Status | Schema | |---|---|---| | garage.rajsingh.info/v1beta2 | Current (storage version, recommended) | Tier-based: spec.storage and/or spec.gateway | | garage.rajsingh.info/v1beta1 | Deprecated, still served | Legacy flat schema: spec.replicas, spec.gateway: bool |

A conversion webhook handles reads and writes in both directions, so existing v1beta1 manifests continue to work unchanged. The controller operates on v1beta2 internally. New clusters should be written as v1beta2.

kubectl scale is supported for an Auto-managed default storage group on both versions: the scale subresource targets .spec.storage.replicas on v1beta2 and .spec.replicas on v1beta1. A gateway-only v1beta1 view retains its historical gateway Scale behavior when clients explicitly target the v1beta1 resource; the preferred v1beta2 endpoint and Manual shapes do not expose a controllable scalable group. A v1beta2 CR that declares both storage and gateway has no faithful v1beta1 form; the conversion webhook returns only the storage tier when read as v1beta1 and marks the v1beta2-only gateway payload. spec.storage.nodeLocalPools is also v1beta2-only. A reserved conversion payload preserves it through a v1beta1 read/write round trip, but v1beta1 clients cannot edit it. Tools that manage either unified tiers or node-local pools must use v1beta2.

Quick Start

First, create an admin token secret for the operator to manage Garage resources:

kubectl create secret generic garage-admin-token \
  --from-literal=admin-token=$(openssl rand -hex 32)

Create a unified 3-storage / 2-gateway Garage cluster (full example):

apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
  name: garage
spec:
  zone: us-east-1
  replication:
    factor: 3
  storage:
    replicas: 3
    metadata:
      size: 10Gi
    data:
      size: 100Gi
  gateway:
    replicas: 2
  network:
    rpcBindPort: 3901
    service:
      type: ClusterIP
  admin:
    adminTokenSecretRef:
      name: garage-admin-token
      key: admin-token

spec.gateway is optional — omit it for a storage-only cluster. Existing v1beta1 manifests (spec.replicas, spec.gateway: bool) are still accepted; the conversion webhook rewrites them to the tier-based shape on read.

Wait for the cluster to be ready:

kubectl wait --for=condition=Ready garagecluster/garage --timeout=300s

Create a bucket:

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageBucket
metadata:
  name: my-bucket
spec:
  clusterRef:
    name: garage
  quotas:
    maxSize: 10Gi

Create access credentials:

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageKey
metadata:
  name: my-key
spec:
  clusterRef:
    name: garage
  bucketPermissions:
    - bucketRef:
        name: my-bucket
      read: true
      write: true

Or grant access to all buckets in the cluster — useful for admin tools, monitoring, or mountpoint-s3 workloads that span multiple buckets:

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageKey
metadata:
  name: admin-key
spec:
  clusterRef:
    name: garage
  allBuckets:
    read: true
    write: true
    owner: true

Per-bucket overrides layer on top of allBuckets, so you can combine cluster-wide read with owner on a specific bucket:

allBuckets:
    read: true
  bucketPermissions:
    - bucketRef:
        name: metrics-bucket
      owner: true

Import existing credentials from an inline spec or a Kubernetes secret:

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageKey
metadata:
  name: imported-key
spec:
  clusterRef:
    name: garage
  importKey:
    # Garage v2.3+: 8+ chars of [A-Za-z0-9-_.] and a 16+ char graphic-ASCII secret.
    # Garage v2.0-v2.2: "GK" + 24 hex characters and a 64-character hex secret.
    accessKeyId: "GK0123456789abcdef01234567"
    secretAccessKey: "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"

Or reference an existing secret — use accessKeyIdKey/secretAccessKeyKey to specify which keys to read from the source secret (defaults to access-key-id/secret-access-key):

importKey:
    secretRef:
      name: my-existing-creds
    accessKeyIdKey: AWS_ACCESS_KEY_ID
    secretAccessKeyKey: AWS_SECRET_ACCESS_KEY

Secret Template

By default the generated secret includes access-key-id, secret-access-key, endpoint, host, scheme, and region. Use secretTemplate to customize what gets included and how keys are named:

secretTemplate:
  accessKeyIdKey: AWS_ACCESS_KEY_ID
  secretAccessKeyKey: AWS_SECRET_ACCESS_KEY
  endpointKey: AWS_ENDPOINT_URL_S3
  regionKey: AWS_REGION
  includeEndpoint: false   # omit endpoint/host/scheme
  includeRegion: false     # omit region
  includeCredentialsFile: true
  credentialsFileKey: credentials
  credentialsFileProfile: default

This is useful when mounting the secret directly as environment variables with envFrom — only the keys your app expects will be present. When includeCredentialsFile is enabled, the selected key contains an AWS shared credentials file using credentialsFileProfile (default default). The file contains only aws_access_key_id and aws_secret_access_key; region and endpoint remain in their separate Secret keys.

Get S3 credentials:

kubectl get secret my-key -o jsonpath='{.data.access-key-id}' | base64 -d && echo
kubectl get secret my-key -o jsonpath='{.data.secret-access-key}' | base64 -d && echo
kubectl get secret my-key -o jsonpath='{.data.endpoint}' | base64 -d && echo

Gateway Tier

spec.gateway runs S3/Admin proxies that store no object blocks (data dir is EmptyDir). Its workload shape depends on the topology:

  • Unified cluster (gateway alongside spec.storage): the gateway tier is reconciled as one per-pod GarageNode (-gateway-N, gateway: true) — symmetric with the storage tier — each owning a single-replica StatefulSet with a small persistent metadata PVC (default 1Gi). Its StatefulSet leaves the Kubernetes PVC-retention policy unset, so the default is Retain on scale-down and deletion.
  • Edge gateway (gateway-only CR + connectTo): the tier stays a single cluster-level StatefulSet (-gateway) because its layout lives on a remote storage cluster. This StatefulSet explicitly uses Delete/Delete PVC retention.
gateway.metadata.type: EmptyDir is an explicit ephemeral-identity option for either managed shape. In Manual unified mode, configure metadata on each user-owned gateway GarageNode; the webhook rejects the unused cluster-level field. Gateway metadata supports the ordinary size, class, access-mode, selector, label, and annotation controls. The selector applies only when a new claim is created in either managed shape and requires a compatible pre-provisioned PV; paths and volumeClaimTemplateSpec are rejected because arbitrary PVC sources can clone or misbind the identity-bearing node_key. To change an edge gateway's metadata source or PVC template, first scale spec.gateway.replicas to zero and wait for its capacity-less roles to retire. This prevents an accepted edit from silently leaving an immutable StatefulSet claim template unchanged.

Gateway pods participate in the cluster layout with capacity: null (matching upstream garage layout assign --gateway). This is required: Garage's S3 sig-auth path uses key_table.get_local() — only nodes in layout.all_nodes() receive FullReplication writes for key_table / bucket_table / admin_token_table. A gateway outside the layout therefore lacks the local authentication record and returns 403 Forbidden: No such key; Garage v2.3.0 does not fall back to a quorum read here. The capacity: null role keeps authentication local and available without an RPC to the storage tier. Scale-downs are tombstone-cleaned (see Gateway tombstone cleanup).

A GarageCluster must set at least one of storage, gateway, or connectTo. The webhook also rejects gateway without either storage (unified pattern) or connectTo (edge pattern). See the gateway examples for more.

Unified cluster (storage + local gateways)

Most common: one CR declares both tiers in the same namespace. Gateway pods talk to the storage tier over the in-cluster RPC service. In Auto mode the operator generates one gateway GarageNode per replica (-gateway-N, gateway: true) alongside the storage tier's -storage-N nodes — both show up in kubectl get gn. Each gateway node gets a capacity: null layout role so key/bucket auth resolves locally. They are operator-owned and are handed off to you on an Auto→Manual flip.

apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
  name: garage
spec:
  zone: us-east-1
  replication:
    factor: 3
  storage:
    replicas: 3
    metadata:
      size: 10Gi
    data:
      size: 100Gi
  gateway:
    replicas: 4
    resources:
      requests:
        cpu: 50m
        memory: 128Mi
  admin:
    adminTokenSecretRef:
      name: garage-admin-token
      key: admin-token

Edge gateway (gateway-only, connects to a remote storage cluster)

For gateways in a different K8s cluster, an external NAS, or a bare-metal Garage instance — omit spec.storage and use connectTo:

apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
  name: garage-edge
spec:
  replication:
    factor: 3        # must match the storage cluster
  gateway:
    # This edge shape has one shared config. Use one replica per independently
    # routed edge identity; use separate edge resources for multiple routes.
    replicas: 1
    # Tells the remote cluster how to dial back to this gateway for bidirectional
    # peering and remote visibility.
    rpcPublicAddr: "edge-gateway.tailnet.example:3901"
  connectTo:
    rpcSecretRef:
      name: garage-rpc-secret
      key: rpc-secret
    adminApiEndpoint: "http://garage-primary.tailnet.example:3903"
    adminTokenSecretRef:
      name: storage-admin-token
      key: admin-token
  admin:
    adminTokenSecretRef:
      name: gateway-admin-token
      key: admin-token
  publicEndpoint:
    type: NodePort
    nodePort:
      basePort: 30901
      externalAddresses:
        - "edge-node1.example.com"
        - "edge-node2.example.com"

Or reference a storage GarageCluster in the same namespace via connectTo.clusterRef.name. The operator opens RPC in both directions (gateway -> external and external -> gateway) when a reverse route is configured; without one, an edge gateway may intentionally run forward-only and the remote site cannot dial or expose that gateway identity. The operator re-establishes configured links on drift; see the gateway sample manifests for complete examples.

Management handle (no owned workload)

A connectTo-only GarageCluster manages buckets, keys, permissions, and layout on an existing Garage deployment without adopting its pods or volumes:

apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
  name: existing-garage
spec:
  connectTo:
    adminApiEndpoint: http://garage.garage.svc:3903
    adminTokenSecretRef:
      name: garage-admin
      key: admin-token
    # Optional: required only to derive new GarageKey material deterministically.
    rpcSecretRef:
      name: garage-rpc
      key: rpc-secret

The operator creates no Garage workload for this shape. When rpcSecretRef is present (or connectTo.clusterRef inherits one), it copies the exact value into an immutable, handle-owned snapshot before reporting ManagementHandleReady. An Admin-only handle needs no RPC secret; imported keys continue to work, and a first RPC source may be attached later, but that source and value cannot then be rotated in place.

Workload differences

| Aspect | Storage tier | Gateway tier (unified) | Gateway tier (edge) | |---|---|---|---| | Workload | N × StatefulSets (one per GarageNode, replicas: 1) | N × StatefulSets (one per gateway GarageNode, replicas: 1) | StatefulSet (-gateway) | | Node CRs | one GarageNode per replica (Auto: operator-owned -storage-N; Manual: user-owned) | one GarageNode per replica (-gateway-N, gateway: true; operator-owned in Auto) | none | | Metadata volume | PVC (per node) | PVC (per node, default 1Gi), or explicit EmptyDir | PVC (default 1Gi), or explicit EmptyDir | | Data volume | PVC (per node) | EmptyDir | EmptyDir | | Pod naming | -0 | -gateway-N-0 | -gateway-0, -gateway-1, … | | Node identity | persists (metadata PVC) | persists (metadata PVC) | persists (metadata PVC) | | Layout owner | per-GarageNode controller (local) | per-GarageNode controller (local), capacity null | remote storage cluster (gateway-connection path) | | Stale-layout cleanup | finalizer on CR deletion | per-node GarageNode finalizer; cluster reaper skips live-claimed roles | operator tombstone-reaps on scale-down |

An externally-routable RPC address is required for bidirectional edge peering and for the remote storage cluster to include the gateway identity in its reachable node view. If you intentionally need only forward connectivity (the gateway can reach the remote cluster, but the remote cluster cannot dial the gateway), omit the address. Garage then advertises the pod IP, the reverse ConnectNode cannot succeed, and the validating webhook emits an admission warning. For a data-less gateway this is a supported forward-only mode; GatewayConnected=True can still mean healthy forward-only connectivity. Set an address when reverse dialing or remote visibility is required. The operator checks three fields, in priority order:

  1. spec.gateway.rpcPublicAddr — preferred for an edge gateway (it has no storage tier to inherit from).
  2. spec.network.rpcPublicAddr.
  3. spec.publicEndpoint — the operator derives the address from the Kubernetes service status.
The validating webhook emits an admission warning when an edge gateway sets connectTo but none of these.

For a single edge identity with bidirectional peering, use publicEndpoint.type: LoadBalancer without loadBalancer.perNode; the operator creates one -rpc LoadBalancer service and derives one rpc_public_addr from it. This is the simplest setup when your infrastructure provides a global/shared load balancer address that routes RPC traffic to that one-replica edge gateway.

For per-pod LoadBalancer services, set publicEndpoint.type: LoadBalancer and publicEndpoint.loadBalancer.perNode: true; the operator creates -0-rpc, -1-rpc, etc. For an edge gateway (single cluster-level StatefulSet sharing one ConfigMap) the operator does not write distinct per-pod rpc_public_addr values into Garage's config; the per-node service addresses are used only when asking the external cluster to connect back to each gateway node. A shared edge config/address is therefore safe only for one independently routed identity. Use separate one-replica edge resources for multiple routes, or use unified gateway GarageNode resources when each identity needs its own advertised address.

The operator establishes connectivity in both directions: gateway → external nodes and external cluster → gateway nodes. It also actively monitors the connection and re-establishes it if Garage marks a peer as unreachable.

Note: bootstrapPeers is also accepted for one-shot bootstrapping when you know the node ID in advance, but adminApiEndpoint is preferred — it works without knowing node IDs upfront and keeps the connection stable across restarts.

Gateway tombstone cleanup

When a gateway scales down, its old capacity: null layout entries must be removed or they inflate the node count that consistent-mode metadata writes (GarageKey/bucket) need for quorum. On each reconcile the operator lists tier:gateway layout entries and cross-references them with the live gateway pods and the node IDs claimed by live operator-owned gateway GarageNodes — a role claimed by an existing GarageNode is never removed, so the cluster reaper never fights the per-node finalizer during a brief pod restart.

Removal is governed by spec.layoutManagement.autoApply:

  • autoApply: true — stale entries are removed and the new layout is applied, then normal Garage history convergence is observed. The operator never runs the cluster-wide skip-dead-nodes recovery automatically.
  • autoApply: false (default) — exact pending IDs are surfaced on status.pendingGatewayTombstones and the GatewayTombstones condition, but are not staged. Remove those exact roles with the Garage CLI or enable autoApply. garage.rajsingh.info/force-layout-apply does not approve tombstones.

Node-local pools (DaemonSet-backed)

Add spec.storage.nodeLocalPools to run node-local, HostPath-backed storage alongside the existing default operator-managed PVC group or hand-managed SMB/PVC GarageNodes. This is the operator's equivalent of the upstream Helm chart's DaemonSet deployment: each entry is realized as one operator-managed DaemonSet running one Garage pod, one generated GarageNode identity, and one Garage layout role per selected Kubernetes Node. The DaemonSet is the current implementation of the node-local contract, not a selectable workload kind, so the field is named for the storage it provides rather than the workload it happens to use. Each named pool has its own capacity, HostPaths, selector, Pod template, and optional per-node RPC address template. Explicitly declaring spec.storage.nodeLocalPools enables the node-local pool API. Node-local pools require Kubernetes 1.27+ for the Pod scheduling-gate safety fence; other cluster shapes retain the chart's Kubernetes 1.25 minimum. The controller verifies the server version, performs a server-side DaemonSet dry run, and requires a real gated probe Pod to receive PodScheduled=False with reason SchedulingGated from kube-scheduler before creating a pool workload or changing Node activation. Successful evidence is namespace-scoped, cached for at most 30 seconds from kube-scheduler's condition timestamp, and pinned only for the reconciliation that began from that proof. The binary and every shipped install enable leader election; silently disabling it is rejected because layout mutation is process-local.

Node-local pool Pods mount HostPaths. Kubernetes Pod Security Admission's Baseline and Restricted policies prohibit HostPath volumes, so each workload namespace that contains a pool needs pod-security.kubernetes.io/enforce=privileged or an equivalent organization-specific exception. This is not required for the operator namespace itself. Grant GarageCluster write access narrowly because a pool Pod template controls access to paths on selected Nodes.

Across all node-local entries, selectors may choose at most 255 Nodes. Garage's global layout limit is 256 positive-capacity roles across every storage backing and federated site. At 255 live roles the controller admits a new node-local identity only when it can prove a retiring generated member in its Kubernetes control plane. That check does not reserve role 256 against independently operated federated sites: another writer can consume the slot and make the node-local assignment wait for a committed removal. Federated topology changes must therefore be externally serialized.

Parent status keeps full per-member detail on label-addressable generated GarageNode resources. Node-local, storage rollout/drain, and repeated-member health condition messages are capped at 4096 bytes with five inventory examples; diagnostic layout history is capped at 64 entries, and a crash-safe rollout records at most 32 retired workload-controller UIDs. A 256-role, eight-resync-worker drain projection measures 365085 bytes (about 357 KiB); a conservative coexisting projection with all 26 GarageCluster conditions at maximum message size measures 966767 bytes (about 944 KiB) and is regression-capped at 1 MiB. These are explicit release-envelope projections, not permission to increase Kubernetes object limits.

Pool membership is drain-safe. The operator translates the user selector into private activation labels, serializes every same-cluster layout writer behind one coordinator, waits for new identities to enter the committed layout, drains retired GarageNodes one at a time while their pods remain online, and removes activation only after finalization. A Kubernetes Node cannot move directly between pools; unselect and fully drain it before selecting the new pool.

Every pool Pod starts scheduler-gated. Only the exact current DaemonSet UID is allowed through after live Node activation, HostPath claim, and competing-Pod checks, so a late Pod from a retired DaemonSet cannot mount the same local disk.

Committed pool members also carry an internal, stable Garage-ID pin on their Kubernetes Node. During a cold or same-name cluster recovery, all exact already-committed roles may restart together, but each child must rediscover that pinned identity from its own Pod and match the operator-tagged committed role before any status or layout write. A wiped or swapped HostPath therefore fails closed instead of enrolling a second identity.

Image, config, and pod-template changes use parent-controlled OnDelete rollout: one pod across the whole GarageCluster is replaced, its identity is rediscovered from the exact replacement pod, and cluster health/layout history must settle before another pod is stopped.

Storage deletion is prepared before Kubernetes DELETE. The source process stays online while Garage removes its role and exact source-plus-destination repair/resync evidence reaches a terminal quiet window; admission then accepts only that exact completed actor. Selector scale-down does this automatically. Direct GarageNode deletion and federated-site retirement use the documented garage.rajsingh.info/drain=true annotate, wait, then delete workflow. Delete an individual GarageNode with default/background propagation; direct foreground deletion is rejected because it could reap the identity-bearing pod before finalizer convergence. A parent GarageCluster may foreground-cascade its children only after its terminal Drain handoff, or as part of explicit whole-store Destroy cleanup.

Pools use the operator-wide Garage v2.0.0+ Admin API v2 floor; v2.4.1 is the tested default. They also require a cluster-scoped install, enabled validating and conversion webhooks, and an Admin API token. One workload-owning GarageCluster is one Garage store/site lifecycle and ownership boundary, usually one physical site. Its members may share one static zone or derive multiple actual failure-domain zones through zoneFrom. Express SMB, local disks, and different local capacities as node-local pools or ordinary GarageNodes inside it. Positive-capacity removals are a separate topology-only generation after configuration rollout has converged. See node-local storage guide, the mixed-storage sample, and the design.

Failure Domains Inside One Cluster (zoneFrom)

spec.zone assigns one Garage zone to the whole cluster, which leaves replication.zoneRedundancyMode with nothing to act on: upstream computes Maximum as min(distinct zones, replication factor), so a single-zone cluster has an effective redundancy of 1. spec.zoneFrom derives each storage node's zone from a label on the Kubernetes Node its pod is scheduled to, so one cluster can express racks, power circuits, or per-node domains without splitting into a federation.

spec:
  zone: site-a                      # fallback, still required
  zoneFrom:
    nodeLabel: topology.kubernetes.io/zone
  replication:
    factor: 3
    zoneRedundancyMode: Maximum     # now meaningful — 3 copies in 3 domains
  storage:
    replicas: 6

Use kubernetes.io/hostname for per-node domains, or a custom label such as example.com/rack for physical racks.

The zone depends on where the pod landed, so it is resolved after scheduling and re-checked on every reconcile. spec.zone is the initial fallback before the Pod is scheduled and whenever the readable Kubernetes Node does not carry the configured label. Once a node has reported an effective status.zone, a transient Pod replacement gap retains that last proven value; this prevents an ordinary rollout from flipping the Garage layout to the fallback zone and back. Failure to read a required Kubernetes Node is different: the resource reports Failed and the operator does not silently mutate the layout using spec.zone. A successfully reconciled node therefore always has a proven zone. The effective value is reported as status.zone on each GarageNode and shown in the ZONE column:

kubectl get garagenodes

NAME CLUSTER ZONE CAPACITY GATEWAY CONNECTED INLAYOUT AGE

garage-storage-0 garage rack-a 500Gi false true true 5m

garage-storage-1 garage rack-b 500Gi false true true 5m

If a pod moves to a Kubernetes Node in a different domain, the layout is updated to match — Garage minimizes the resulting reassignment rather than reshuffling everything. Nodes whose PVCs pin them to a machine will not move at all.

Scope and caveats:

  • Operator-managed members. Cluster-level zoneFrom applies to generated default-group GarageNodes and to nodeLocalPools, including when storage.layoutPolicy: Manual disables only the default group. Set zoneFrom directly on each user-owned Manual GarageNode.
  • Storage tier only. It is deliberately not applied to gateway nodes: Garage counts gateway zones toward the Maximum redundancy target but satisfies that target from storage nodes only, so per-node gateway zones can make every layout apply fail.
  • zoneRedundancyMode: AtLeast(n) requires the label to resolve to at least n distinct values across scheduled storage pods, otherwise layout apply is rejected upstream. The webhook warns about this combination.
  • Requires the cluster-scoped install (the default). A namespace-scoped install cannot read Nodes, so affected GarageClusters and GarageNodes report Failed instead of silently falling back to spec.zone.

Manual Node Layout (GarageNode)

By default, GarageCluster uses layoutPolicy: Auto — the operator generates one operator-owned GarageNode per storage replica (named -storage-N), and each GarageNode controller drives its own single-replica StatefulSet. For fine-grained control over individual nodes (per-node zone, capacity, tags, storage class, RPC address, external nodes), set layoutPolicy: Manual and create GarageNode resources directly.

Each GarageNode creates a single-replica StatefulSet and manages that node's layout entry (zone, capacity, tags). Flipping a cluster from Auto → Manual is a one-way hand-off: the operator drops its controllerRef on each -storage-N GarageNode and the user inherits them. The reverse (Manual → Auto) is rejected by the validating webhook.

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageNode
metadata:
  name: storage-node-a
spec:
  clusterRef:
    name: garage
  zone: zone-a
  capacity: 500Gi
  tags: ["ssd", "high-performance"]
  storage:
    metadata:
      size: 10Gi
    data:
      size: 500Gi
      storageClassName: fast-ssd

Gateway Nodes

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageNode
metadata:
  name: gateway-node
spec:
  clusterRef:
    name: garage
  zone: zone-a
  gateway: true
  storage:
    metadata:
      size: 1Gi

External Nodes

For nodes running outside Kubernetes (bare-metal, NAS, other clusters):

apiVersion: garage.rajsingh.info/v1beta1
kind: GarageNode
metadata:
  name: external-node
spec:
  clusterRef:
    name: garage
  nodeId: "563e1ac825ee3323aa441e72c26d1030d6d4414aeb3dd25287c531e7fc2bc95d"
  zone: dc-1
  capacity: 1Ti
  external:
    address: nas.local
    port: 3901

External nodes require nodeId (64-hex-char Ed25519 public key). No StatefulSet is created — the operator only manages the layout entry.

Per-Node Overrides

GarageNode supports overriding cluster defaults: image, imageRepository, resources, nodeSelector, tolerations, affinity, podAnnotations, podLabels, priorityClassName, imagePullPolicy, imagePullSecrets, serviceAccountName, securityContext, containerSecurityContext, topologySpreadConstraints, env, envFrom, and logging, plus per-node network.rpcPublicAddr, publicEndpoint, and storage (fsync, snapshots, dataPaths). A node with any of these gets its own -config ConfigMap instead of the shared cluster config.

Per-Node Maintenance

Set spec.maintenance.suspended: true on a single GarageNode to pause reconciliation of just that node's StatefulSet, ConfigMap, Service, and layout entry — useful for PVC-level work (StorageClass migration, longhorn engine upgrade, disk swap) without the controller fighting you. A Suspended status condition is set while paused, and the finalizer/delete path still runs so a suspended node can be deleted.

Status

kubectl get garagenodes

NAME CLUSTER ZONE CAPACITY GATEWAY CONNECTED INLAYOUT AGE

storage-node-a garage zone-a 500Gi false true true 5m

The controller auto-discovers node IDs from pods, reconciles layout drift (zone/capacity/tags), and handles node removal with replication-safe finalization.

Scaling

GarageCluster supports the Kubernetes scale subresource for the Auto-managed default storage group, enabling kubectl scale and compatibility with autoscalers such as VPA and HPA for that workload.

kubectl scale garagecluster garage --replicas=5

The scale subresource targets .spec.storage.replicas on v1beta2 and .spec.replicas on v1beta1. It controls only the Auto-managed default StatefulSet/PVC group. Node-local pool cardinality follows each storage.nodeLocalPools[].selector; gateway-tier replicas are not exposed through v1beta2 /scale, so adjust spec.gateway.replicas directly. A gateway-only v1beta1 object retains its historical gateway Scale mapping. Dedicated status.scaleReplicas/status.scaleSelector report the actual non-terminating Pods and exact selector for the controllable workload; aggregate status remains separate. Manual storage has no scalable default group—its ordinary GarageNodes are individually owned resources—so HPA/VPA and kubectl scale are unsupported for that shape. A separate fail-closed admission handler runs the full GarageCluster topology validation for /scale, including active drain and prepared scale-down gates.

Because v1beta2 is the preferred discovery version, legacy edge-gateway scaling must name the v1beta1 resource explicitly:

kubectl scale garageclusters.v1beta1.garage.rajsingh.info edge --replicas=5

PVC Retention Policy

By default, PVCs created by a GarageCluster's StatefulSet are not deleted when the cluster is deleted or scaled down. This is intentional: Garage stores your data in those volumes, and automatic deletion would be irreversible.

Storage-member behavior is controlled by spec.storage.pvcRetentionPolicy:

| Field | Value | Behavior | |-------|-------|----------| | whenDeleted | Retain (default) | PVCs survive GarageCluster deletion — manual cleanup required | | whenDeleted | Delete | PVCs are deleted automatically when the GarageCluster is deleted | | whenScaled | Retain (default) | PVCs for scaled-down pods are kept (allows scaling back up) | | whenScaled | Delete | PVCs for removed replicas are deleted on scale-down |

For dev/test clusters where you want automatic cleanup:

spec:
  storage:
    pvcRetentionPolicy:
      whenDeleted: Delete
      whenScaled: Delete

Requires Kubernetes 1.23+. For production clusters, leave this unset (defaults to Retain) or set whenScaled: Delete only if you're confident scaled-down nodes won't need their data again.

Gateway metadata has a separate spec.gateway.pvcRetentionPolicy. An Auto unified gateway member owns a single-replica StatefulSet and defaults to Kubernetes Retain/Retain; an edge gateway keeps its released cluster-level StatefulSet default of Delete/Delete. An explicit gateway policy applies to both managed shapes and never changes storage-member claims. A v1beta1 edge gateway's released spec.storage.pvcRetentionPolicy is losslessly projected to this field by conversion.

spec:
  gateway:
    pvcRetentionPolicy:
      whenDeleted: Retain
      whenScaled: Retain

Choose Delete only when losing that gateway's persisted node_key after its capacity-less layout role is retired is intentional. Gateway claims contain no object blocks, but deleting metadata still creates a different Garage identity on the next start.

Note (Auto mode): automatic storage and gateway topology changes wait until every earlier Garage layout version has left Draining. The initial storage bootstrap creates the required members together because no Admin API exists yet; every later scale-up or gateway addition admits one per-node GarageNode at a time. Scale-down likewise drains one member at a time and does not begin the next removal until the prior finalizer completes. StorageTopologyReady=False reports AddingMembers, DrainingMember, or WaitingForLayoutSync, and cluster Ready=False; StorageScaleDownBlocked=True is reserved for a scale-down that would violate replication.factor. A removed node's single-replica StatefulSet is deleted, so its PVCs are governed by whenDeleted, not whenScaled. To reclaim those volumes, set whenDeleted: Delete.

GarageNode Replacement Cycles

garage.rajsingh.info/cycle=true is a narrow add-before-remove automation for an established, positive-capacity, StatefulSet-backed GarageNode. Add the annotation only after the source's exact Pod and 64-hex Garage node ID are Ready, connected, committed to a settled layout, and observed at the current generation:

kubectl -n garage-operator-system annotate garagenode garage-storage-a \
  garage.rajsingh.info/cycle=true

The operator creates one sibling with a fresh Garage identity and fresh claim names from the same repeatable PVC templates (or the same explicit EmptyDir profile), waits for that exact sibling Pod to become Ready and enter the settled layout, then runs the ordinary cluster-wide drain and block-resync proof before promoting it. A PVC selector is repeated on the new claim, so static-PV users must provision distinct replacement PVs in advance.

Automatic cycle deliberately rejects existingClaim, gateways, external processes, and node-local-pool members. It never infers replacement hardware or a StorageClass from a bound claim, and it never reuses, clones, snapshots, or deletes source claims. For SMB, Ceph, manually bound PVCs, exceptional members, or a different disk profile, explicitly create a second GarageNode with distinct metadata and data storage, wait for it to synchronize, then drain and delete the old identity. Node-local membership changes use spec.storage.nodeLocalPools[].selector and the pool retirement state machine instead of this annotation.

Progress is durable in status.cyclePhase and the Cycling condition. Removing the request before the source enters Draining is only a cancellation request: an already-created sibling must first be explicitly drained and deleted. Once the source has entered Draining, the annotation and transaction are one-way; the controller fails closed unless the exact persisted sibling identity appears in the terminal drain proof.

Multi-HDD Storage

Garage supports striping a node's data across multiple disks. To use it, set spec.storage.data.paths[] instead of spec.storage.data.size — the operator emits one PVC + one volumeMount per path, and renders the matching data_dir TOML array.

spec:
  storage:
    replicas: 3
    metadata:
      size: 10Gi
    data:
      paths:
        - path: /data/data0
          volume:
            size: 1Ti
            storageClassName: fast-ssd
        - path: /data/data1
          volume:
            size: 4Ti
            storageClassName: bulk-hdd
        - path: /mnt/archive
          readOnly: true   # legacy disk, read-only mount, no capacity

Each storage replica is its own single-replica StatefulSet -storage-, so PVCs follow the