Garage Kubernetes Operator
S3-Compatible Object Storage on Kubernetes
Read the operator documentation · Releases · Support
A Kubernetes operator for Garage - distributed, self-hosted object storage with multi-cluster federation.
- Declarative cluster lifecycle — StatefulSet, config, and layout managed via CRDs
- Unified storage + gateway tiers in one CR (v1beta2) — combine durable storage pods and persistent-identity S3 gateways in a single
GarageCluster - Node-local pools — bind Garage identities to selected Kubernetes Nodes and HostPath disks, including multi-disk layouts
- Bucket & key management — create buckets, quotas, and S3 credentials with kubectl
- Multi-cluster federation — span storage across Kubernetes clusters with automatic node discovery
- Persistent-identity gateway pods — StatefulSet with a small metadata PVC; gateway pods keep the same Garage node identity across restarts and participate in the cluster layout with
capacity: null(matching upstreamgarage layout assign --gateway) - Scale subresource —
kubectl scaleand autoscaler support for the Auto-managed default storage group (and v1beta1 edge gateways) - COSI driver — optional Kubernetes-native object storage provisioning
Custom Resources
| CRD | Description |
|-----|-------------|
| GarageCluster | Deploys and manages a Garage cluster (storage and/or gateway tiers) |
| GarageBucket | Creates buckets with quotas and website hosting |
| GarageKey | Provisions S3 access keys with per-bucket permissions |
| GarageNode | Fine-grained node layout control (zone, capacity, tags) |
| GarageAdminToken | Creates static Admin bootstrap material in a namespace-local Secret |
| GarageReferenceGrant | Grants selected cross-namespace access to clusters, buckets, and keys |
Install
Requires Kubernetes 1.25+, or **1.27+ if you use node-local pools**, which depend on Pod scheduling gates for their activation fence.
The Helm chart enables admission and conversion webhooks by default, so install
cert-manager first. Disabling webhooks is limited to local development or
simple v1beta2-only installs that neither use nodeLocalPools nor rely on
admission-protected storage deletion. It removes those safety checks and all
v1beta1 conversion support; node-local pools and controller-managed persistent
claims do not support that mode. EmptyDir remains fully supported there;
explicit existingClaim volumes can be mounted, but their PVC-backed rollout
and recovery paths remain fenced until admission is enabled. The
webhooks also reserve managed PVC finalizer removal to the operator service
account, preventing namespace users with PVC update rights from reopening a
same-name claim replacement race before StatefulSet ownership is established.
helm install garage-operator oci://ghcr.io/rajsinghtech/charts/garage-operator \
--namespace garage-operator-system \
--create-namespace
helm install garage-operator oci://ghcr.io/rajsinghtech/charts/garage-operator \
--namespace garage-operator-system \
--create-namespace \
--set webhooks.enabled=false
Verifying release artifacts
Released container images and Helm charts are signed with cosign keyless signing (the GitHub Actions OIDC identity — no long-lived keys), and carry SLSA build provenance. The image additionally carries an SPDX SBOM. All three are stored in GHCR as OCI referrers of the artifact digest.
IMAGE=ghcr.io/rajsinghtech/garage-operator:v0.8.0
Signature
cosign verify "$IMAGE" \
--certificate-identity-regexp '^https://github.com/rajsinghtech/garage-operator/\.github/workflows/docker\.yml@refs/' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
Provenance and SBOM
gh attestation verify "oci://$IMAGE" --repo rajsinghtech/garage-operator
cosign download attestation "$IMAGE" --predicate-type https://spdx.dev/Document/v2.3
The Helm chart is signed the same way (--certificate-identity-regexp ending in helm\.yml@refs/), and dist/install.yaml attached to each GitHub release has a provenance attestation verifiable with gh attestation verify install.yaml --repo rajsinghtech/garage-operator.
Under a policy controller, pin by digest and require the signature — e.g. Kyverno verifyImages with keyless.issuer: https://token.actions.githubusercontent.com and the subject regexp above.
Garage Version Compatibility
The Garage version is yours to choose — GarageCluster.spec.image, GarageNode.spec.image, or the chart-wide defaultGarageImage. The chart's appVersion tracks the operator, not Garage.
| Operator | Garage minimum | Garage tested in CI | Notes |
|---|---|---|---|
| 0.8.x | v2.0.0 | v2.4.1 (default, all suites), v2.0.0 (floor lane); nightly main-v2 canary | Admin API v2; node-local pools require Kubernetes 1.27+ |
| 0.7.x | v2.0.0 | v2.4.0, v2.2.0 | Admin API v2; node-local pools require Kubernetes 1.27+ |
| 0.6.x | v2.0.0 | v2.3.0, v2.2.0 | Admin API v2 only |
dxflrs/garage:v2.4.1@sha256:9c96caa2612d3411acc5b0e6701fb238dbfba33e533a6d7d3d811a4b12d0d020 is the built-in default when spec.image is unset, so default deployments run the multi-platform image index that the Ginkgo suite and the topology suites (multi-cluster, external gateway, IPv6, single-cluster) are pinned to. Two further lanes back the "v2.x range" claim: a floor lane runs the core bucket and key path on the pinned Garage v2.0.0 index (the oldest supported release; v2.0.0 is the only v2.0.x release), and a nightly canary runs the same path on a build of Garage's main-v2 branch (.github/workflows/garage-canary.yml). One lane enables Garage's native kubernetes_discovery with namespaced RBAC. Any other Garage image, including every other release, is compatible by contract (the Admin API v2) but not individually tested; see the compatibility matrix.
Garage 0.x and 1.x are not supported. The operator drives buckets, keys, layout, and repair exclusively through the /v2/... admin API, which first shipped in Garage v2.0.0. Against an older node every admin call 404s and no cluster will reconcile.
Some fields need a newer Garage than the v2.0.0 floor:
| Field | Requires | Behavior on older Garage |
|---|---|---|
| GarageBucket.spec.lifecycle | v2.3.0 | Older nodes accept the write and drop the field. The operator reads the rules back and sets LifecycleConfigured=False naming this requirement, rather than reporting a success that never took effect. The bucket itself still reconciles. |
| GarageCluster.spec.database.engine: fjall, spec.database.fjallBlockCacheSize | v2.1.0 | Unknown config key, silently ignored; Garage falls back to the default engine |
| GarageCluster.spec.blocks.maxConcurrentReads | v2.1.0 | Silently ignored |
| GarageCluster.spec.blocks.maxConcurrentWritesPerRequest | v2.2.0 | Silently ignored |
Garage's TOML parser ignores unknown keys, so setting a too-new config field degrades to a no-op rather than a crashloop. The operator only emits these keys when you set the corresponding field.
The Garage version each cluster is actually running is reported back on the CR:
kubectl get garagecluster garage -o jsonpath='{.status.buildInfo.version}'
API Versions
GarageCluster is served under two API versions; all other CRDs are v1beta1.
| Version | Status | Schema |
|---|---|---|
| garage.rajsingh.info/v1beta2 | Current (storage version, recommended) | Tier-based: spec.storage and/or spec.gateway |
| garage.rajsingh.info/v1beta1 | Deprecated, still served | Legacy flat schema: spec.replicas, spec.gateway: bool |
A conversion webhook handles reads and writes in both directions, so existing v1beta1 manifests continue to work unchanged. The controller operates on v1beta2 internally. New clusters should be written as v1beta2.
kubectl scale is supported for an Auto-managed default storage group on both
versions: the scale subresource targets .spec.storage.replicas on v1beta2 and
.spec.replicas on v1beta1. A gateway-only v1beta1 view retains its historical
gateway Scale behavior when clients explicitly target the v1beta1 resource;
the preferred v1beta2 endpoint and Manual shapes do not expose a controllable
scalable group. A v1beta2
CR that declares both storage and gateway has no faithful v1beta1 form;
the conversion webhook returns only the storage tier when read as v1beta1 and
marks the v1beta2-only gateway payload. spec.storage.nodeLocalPools is also
v1beta2-only. A reserved conversion payload preserves it through a v1beta1
read/write round trip, but v1beta1 clients cannot edit it. Tools that manage
either unified tiers or node-local pools must use v1beta2.
Quick Start
First, create an admin token secret for the operator to manage Garage resources:
kubectl create secret generic garage-admin-token \
--from-literal=admin-token=$(openssl rand -hex 32)
Create a unified 3-storage / 2-gateway Garage cluster (full example):
apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
name: garage
spec:
zone: us-east-1
replication:
factor: 3
storage:
replicas: 3
metadata:
size: 10Gi
data:
size: 100Gi
gateway:
replicas: 2
network:
rpcBindPort: 3901
service:
type: ClusterIP
admin:
adminTokenSecretRef:
name: garage-admin-token
key: admin-token
spec.gateway is optional — omit it for a storage-only cluster. Existing v1beta1 manifests (spec.replicas, spec.gateway: bool) are still accepted; the conversion webhook rewrites them to the tier-based shape on read.
Wait for the cluster to be ready:
kubectl wait --for=condition=Ready garagecluster/garage --timeout=300s
Create a bucket:
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageBucket
metadata:
name: my-bucket
spec:
clusterRef:
name: garage
quotas:
maxSize: 10Gi
Create access credentials:
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageKey
metadata:
name: my-key
spec:
clusterRef:
name: garage
bucketPermissions:
- bucketRef:
name: my-bucket
read: true
write: true
Or grant access to all buckets in the cluster — useful for admin tools, monitoring, or mountpoint-s3 workloads that span multiple buckets:
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageKey
metadata:
name: admin-key
spec:
clusterRef:
name: garage
allBuckets:
read: true
write: true
owner: true
Per-bucket overrides layer on top of allBuckets, so you can combine cluster-wide read with owner on a specific bucket:
allBuckets:
read: true
bucketPermissions:
- bucketRef:
name: metrics-bucket
owner: true
Import existing credentials from an inline spec or a Kubernetes secret:
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageKey
metadata:
name: imported-key
spec:
clusterRef:
name: garage
importKey:
# Garage v2.3+: 8+ chars of [A-Za-z0-9-_.] and a 16+ char graphic-ASCII secret.
# Garage v2.0-v2.2: "GK" + 24 hex characters and a 64-character hex secret.
accessKeyId: "GK0123456789abcdef01234567"
secretAccessKey: "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
Or reference an existing secret — use accessKeyIdKey/secretAccessKeyKey to specify which keys to read from the source secret (defaults to access-key-id/secret-access-key):
importKey:
secretRef:
name: my-existing-creds
accessKeyIdKey: AWS_ACCESS_KEY_ID
secretAccessKeyKey: AWS_SECRET_ACCESS_KEY
Secret Template
By default the generated secret includes access-key-id, secret-access-key, endpoint, host, scheme, and region. Use secretTemplate to customize what gets included and how keys are named:
secretTemplate:
accessKeyIdKey: AWS_ACCESS_KEY_ID
secretAccessKeyKey: AWS_SECRET_ACCESS_KEY
endpointKey: AWS_ENDPOINT_URL_S3
regionKey: AWS_REGION
includeEndpoint: false # omit endpoint/host/scheme
includeRegion: false # omit region
includeCredentialsFile: true
credentialsFileKey: credentials
credentialsFileProfile: default
This is useful when mounting the secret directly as environment variables with envFrom — only the keys your app expects will be present.
When includeCredentialsFile is enabled, the selected key contains an AWS
shared credentials file using credentialsFileProfile (default default).
The file contains only aws_access_key_id and aws_secret_access_key; region
and endpoint remain in their separate Secret keys.
Get S3 credentials:
kubectl get secret my-key -o jsonpath='{.data.access-key-id}' | base64 -d && echo
kubectl get secret my-key -o jsonpath='{.data.secret-access-key}' | base64 -d && echo
kubectl get secret my-key -o jsonpath='{.data.endpoint}' | base64 -d && echo
Gateway Tier
spec.gateway runs S3/Admin proxies that store no object blocks (data dir is EmptyDir). Its workload shape depends on the topology:
- Unified cluster (gateway alongside
spec.storage): the gateway tier is reconciled as one per-podGarageNode(,-gateway-N gateway: true) — symmetric with the storage tier — each owning a single-replicaStatefulSetwith a small persistent metadata PVC (default 1Gi). Its StatefulSet leaves the Kubernetes PVC-retention policy unset, so the default isRetainon scale-down and deletion. - Edge gateway (gateway-only CR +
connectTo): the tier stays a single cluster-levelStatefulSet() because its layout lives on a remote storage cluster. This StatefulSet explicitly uses-gateway Delete/DeletePVC retention.
gateway.metadata.type: EmptyDir is an explicit ephemeral-identity option for
either managed shape. In Manual unified mode, configure metadata on each
user-owned gateway GarageNode; the webhook rejects the unused cluster-level
field. Gateway metadata supports the ordinary size, class, access-mode,
selector, label, and annotation controls. The selector applies only when a new
claim is created in either managed shape and requires a compatible
pre-provisioned PV; paths and
volumeClaimTemplateSpec are rejected because arbitrary PVC sources can clone
or misbind the identity-bearing node_key. To change an edge gateway's
metadata source or PVC template, first scale
spec.gateway.replicas to zero and wait for its capacity-less roles to retire.
This prevents an accepted edit from silently leaving an immutable StatefulSet
claim template unchanged.
Gateway pods participate in the cluster layout with capacity: null (matching upstream garage layout assign --gateway). This is required: Garage's S3 sig-auth path uses key_table.get_local() — only nodes in layout.all_nodes() receive FullReplication writes for key_table / bucket_table / admin_token_table. A gateway outside the layout therefore lacks the local authentication record and returns 403 Forbidden: No such key; Garage v2.3.0 does not fall back to a quorum read here. The capacity: null role keeps authentication local and available without an RPC to the storage tier. Scale-downs are tombstone-cleaned (see Gateway tombstone cleanup).
A GarageCluster must set at least one of storage, gateway, or connectTo. The webhook also rejects gateway without either storage (unified pattern) or connectTo (edge pattern). See the gateway examples for more.
Unified cluster (storage + local gateways)
Most common: one CR declares both tiers in the same namespace. Gateway pods talk to the storage tier over the in-cluster RPC service. In Auto mode the operator generates one gateway GarageNode per replica (, gateway: true) alongside the storage tier's nodes — both show up in kubectl get gn. Each gateway node gets a capacity: null layout role so key/bucket auth resolves locally. They are operator-owned and are handed off to you on an Auto→Manual flip.
apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
name: garage
spec:
zone: us-east-1
replication:
factor: 3
storage:
replicas: 3
metadata:
size: 10Gi
data:
size: 100Gi
gateway:
replicas: 4
resources:
requests:
cpu: 50m
memory: 128Mi
admin:
adminTokenSecretRef:
name: garage-admin-token
key: admin-token
Edge gateway (gateway-only, connects to a remote storage cluster)
For gateways in a different K8s cluster, an external NAS, or a bare-metal Garage instance — omit spec.storage and use connectTo:
apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
name: garage-edge
spec:
replication:
factor: 3 # must match the storage cluster
gateway:
# This edge shape has one shared config. Use one replica per independently
# routed edge identity; use separate edge resources for multiple routes.
replicas: 1
# Tells the remote cluster how to dial back to this gateway for bidirectional
# peering and remote visibility.
rpcPublicAddr: "edge-gateway.tailnet.example:3901"
connectTo:
rpcSecretRef:
name: garage-rpc-secret
key: rpc-secret
adminApiEndpoint: "http://garage-primary.tailnet.example:3903"
adminTokenSecretRef:
name: storage-admin-token
key: admin-token
admin:
adminTokenSecretRef:
name: gateway-admin-token
key: admin-token
publicEndpoint:
type: NodePort
nodePort:
basePort: 30901
externalAddresses:
- "edge-node1.example.com"
- "edge-node2.example.com"
Or reference a storage GarageCluster in the same namespace via connectTo.clusterRef.name. The operator opens RPC in both directions (gateway -> external and external -> gateway) when a reverse route is configured; without one, an edge gateway may intentionally run forward-only and the remote site cannot dial or expose that gateway identity. The operator re-establishes configured links on drift; see the gateway sample manifests for complete examples.
Management handle (no owned workload)
A connectTo-only GarageCluster manages buckets, keys, permissions, and
layout on an existing Garage deployment without adopting its pods or volumes:
apiVersion: garage.rajsingh.info/v1beta2
kind: GarageCluster
metadata:
name: existing-garage
spec:
connectTo:
adminApiEndpoint: http://garage.garage.svc:3903
adminTokenSecretRef:
name: garage-admin
key: admin-token
# Optional: required only to derive new GarageKey material deterministically.
rpcSecretRef:
name: garage-rpc
key: rpc-secret
The operator creates no Garage workload for this shape. When rpcSecretRef is
present (or connectTo.clusterRef inherits one), it copies the exact value into
an immutable, handle-owned snapshot before reporting ManagementHandleReady.
An Admin-only handle needs no RPC secret; imported keys continue to work, and a
first RPC source may be attached later, but that source and value cannot then be
rotated in place.
Workload differences
| Aspect | Storage tier | Gateway tier (unified) | Gateway tier (edge) |
|---|---|---|---|
| Workload | N × StatefulSets (one per GarageNode, replicas: 1) | N × StatefulSets (one per gateway GarageNode, replicas: 1) | StatefulSet () |
| Node CRs | one GarageNode per replica (Auto: operator-owned ; Manual: user-owned) | one GarageNode per replica (, gateway: true; operator-owned in Auto) | none |
| Metadata volume | PVC (per node) | PVC (per node, default 1Gi), or explicit EmptyDir | PVC (default 1Gi), or explicit EmptyDir |
| Data volume | PVC (per node) | EmptyDir | EmptyDir |
| Pod naming | | | , , … |
| Node identity | persists (metadata PVC) | persists (metadata PVC) | persists (metadata PVC) |
| Layout owner | per-GarageNode controller (local) | per-GarageNode controller (local), capacity null | remote storage cluster (gateway-connection path) |
| Stale-layout cleanup | finalizer on CR deletion | per-node GarageNode finalizer; cluster reaper skips live-claimed roles | operator tombstone-reaps on scale-down |
An externally-routable RPC address is required for bidirectional edge peering and
for the remote storage cluster to include the gateway identity in its reachable
node view. If you intentionally need only forward connectivity (the gateway can
reach the remote cluster, but the remote cluster cannot dial the gateway), omit
the address. Garage then advertises the pod IP, the reverse ConnectNode cannot
succeed, and the validating webhook emits an admission warning. For a data-less
gateway this is a supported forward-only mode; GatewayConnected=True can still
mean healthy forward-only connectivity. Set an address when reverse dialing or
remote visibility is required. The operator checks three fields, in priority order:
spec.gateway.rpcPublicAddr— preferred for an edge gateway (it has no storage tier to inherit from).spec.network.rpcPublicAddr.spec.publicEndpoint— the operator derives the address from the Kubernetes service status.
connectTo but none of these.
For a single edge identity with bidirectional peering, use publicEndpoint.type: LoadBalancer without loadBalancer.perNode; the operator creates one LoadBalancer service and derives one rpc_public_addr from it. This is the simplest setup when your infrastructure provides a global/shared load balancer address that routes RPC traffic to that one-replica edge gateway.
For per-pod LoadBalancer services, set publicEndpoint.type: LoadBalancer and publicEndpoint.loadBalancer.perNode: true; the operator creates , , etc. For an edge gateway (single cluster-level StatefulSet sharing one ConfigMap) the operator does not write distinct per-pod rpc_public_addr values into Garage's config; the per-node service addresses are used only when asking the external cluster to connect back to each gateway node. A shared edge config/address is therefore safe only for one independently routed identity. Use separate one-replica edge resources for multiple routes, or use unified gateway GarageNode resources when each identity needs its own advertised address.
The operator establishes connectivity in both directions: gateway → external nodes and external cluster → gateway nodes. It also actively monitors the connection and re-establishes it if Garage marks a peer as unreachable.
Note:bootstrapPeersis also accepted for one-shot bootstrapping when you know the node ID in advance, butadminApiEndpointis preferred — it works without knowing node IDs upfront and keeps the connection stable across restarts.
Gateway tombstone cleanup
When a gateway scales down, its old capacity: null layout entries must be removed or they inflate the node count that consistent-mode metadata writes (GarageKey/bucket) need for quorum. On each reconcile the operator lists tier:gateway layout entries and cross-references them with the live gateway pods and the node IDs claimed by live operator-owned gateway GarageNodes — a role claimed by an existing GarageNode is never removed, so the cluster reaper never fights the per-node finalizer during a brief pod restart.
Removal is governed by spec.layoutManagement.autoApply:
autoApply: true— stale entries are removed and the new layout is applied, then normal Garage history convergence is observed. The operator never runs the cluster-wideskip-dead-nodesrecovery automatically.autoApply: false(default) — exact pending IDs are surfaced onstatus.pendingGatewayTombstonesand theGatewayTombstonescondition, but are not staged. Remove those exact roles with the Garage CLI or enableautoApply.garage.rajsingh.info/force-layout-applydoes not approve tombstones.
Node-local pools (DaemonSet-backed)
Add spec.storage.nodeLocalPools to run node-local, HostPath-backed storage
alongside the existing default operator-managed PVC group or hand-managed
SMB/PVC GarageNodes. This is the operator's equivalent of the upstream Helm
chart's DaemonSet deployment: each entry is realized as one operator-managed
DaemonSet running one Garage pod, one generated GarageNode identity, and
one Garage layout role per selected Kubernetes Node. The DaemonSet is the
current implementation of the node-local contract, not a selectable workload
kind, so the field is named for the storage it provides rather than the
workload it happens to use.
Each named pool has its own capacity, HostPaths, selector, Pod template, and
optional per-node RPC address template. Explicitly declaring
spec.storage.nodeLocalPools enables the node-local pool API. Node-local
pools require Kubernetes 1.27+ for the
Pod scheduling-gate safety fence; other cluster shapes retain the chart's
Kubernetes 1.25 minimum. The controller verifies the server version, performs a
server-side DaemonSet dry run, and requires a real gated probe Pod to receive
PodScheduled=False with reason SchedulingGated from kube-scheduler before
creating a pool workload or changing Node activation. Successful evidence is
namespace-scoped, cached for at most 30 seconds from kube-scheduler's condition
timestamp, and pinned only for the reconciliation that began from that proof.
The binary and every shipped install enable leader election;
silently disabling it is rejected because layout mutation is process-local.
Node-local pool Pods mount HostPaths. Kubernetes Pod Security Admission's
Baseline and Restricted policies prohibit HostPath volumes, so each workload
namespace that contains a pool needs pod-security.kubernetes.io/enforce=privileged
or an equivalent organization-specific exception. This is not required for the
operator namespace itself. Grant GarageCluster write access narrowly because a
pool Pod template controls access to paths on selected Nodes.
Across all node-local entries, selectors may choose at most 255 Nodes. Garage's global layout limit is 256 positive-capacity roles across every storage backing and federated site. At 255 live roles the controller admits a new node-local identity only when it can prove a retiring generated member in its Kubernetes control plane. That check does not reserve role 256 against independently operated federated sites: another writer can consume the slot and make the node-local assignment wait for a committed removal. Federated topology changes must therefore be externally serialized.
Parent status keeps full per-member detail on label-addressable generated
GarageNode resources. Node-local, storage rollout/drain, and repeated-member
health condition messages are capped at 4096 bytes with five inventory examples;
diagnostic layout history is capped at 64 entries, and a crash-safe
rollout records at most 32 retired workload-controller UIDs. A 256-role,
eight-resync-worker drain projection measures 365085 bytes (about 357 KiB); a
conservative coexisting projection with all 26 GarageCluster conditions at
maximum message size measures 966767 bytes (about 944 KiB) and is
regression-capped at 1 MiB. These are explicit release-envelope projections,
not permission to increase Kubernetes object limits.
Pool membership is drain-safe. The operator translates the user selector into private activation labels, serializes every same-cluster layout writer behind one coordinator, waits for new identities to enter the committed layout, drains retired GarageNodes one at a time while their pods remain online, and removes activation only after finalization. A Kubernetes Node cannot move directly between pools; unselect and fully drain it before selecting the new pool.
Every pool Pod starts scheduler-gated. Only the exact current DaemonSet UID is allowed through after live Node activation, HostPath claim, and competing-Pod checks, so a late Pod from a retired DaemonSet cannot mount the same local disk.
Committed pool members also carry an internal, stable Garage-ID pin on their Kubernetes Node. During a cold or same-name cluster recovery, all exact already-committed roles may restart together, but each child must rediscover that pinned identity from its own Pod and match the operator-tagged committed role before any status or layout write. A wiped or swapped HostPath therefore fails closed instead of enrolling a second identity.
Image, config, and pod-template changes use parent-controlled OnDelete
rollout: one pod across the whole GarageCluster is replaced, its identity is
rediscovered from the exact replacement pod, and cluster health/layout history
must settle before another pod is stopped.
Storage deletion is prepared before Kubernetes DELETE. The source process
stays online while Garage removes its role and exact source-plus-destination
repair/resync evidence reaches a terminal quiet window; admission then accepts
only that exact completed actor. Selector scale-down does this automatically.
Direct GarageNode deletion and federated-site retirement use the documented
garage.rajsingh.info/drain=true annotate, wait, then delete workflow.
Delete an individual GarageNode with default/background propagation; direct
foreground deletion is rejected because it could reap the identity-bearing pod
before finalizer convergence. A parent GarageCluster may foreground-cascade
its children only after its terminal Drain handoff, or as part of explicit
whole-store Destroy cleanup.
Pools use the operator-wide Garage v2.0.0+ Admin API v2 floor; v2.4.1 is the
tested default. They also require a cluster-scoped install, enabled validating
and conversion webhooks, and an Admin API token. One workload-owning
GarageCluster is one Garage store/site lifecycle and ownership boundary,
usually one physical site. Its members may share one static zone or derive
multiple actual failure-domain zones through zoneFrom. Express SMB, local
disks, and different local capacities as node-local pools or ordinary
GarageNodes inside it. Positive-capacity removals are a
separate topology-only generation after configuration rollout has converged.
See
node-local storage guide, the
mixed-storage sample,
and the design.
Failure Domains Inside One Cluster (zoneFrom)
spec.zone assigns one Garage zone to the whole cluster, which leaves replication.zoneRedundancyMode with nothing to act on: upstream computes Maximum as min(distinct zones, replication factor), so a single-zone cluster has an effective redundancy of 1. spec.zoneFrom derives each storage node's zone from a label on the Kubernetes Node its pod is scheduled to, so one cluster can express racks, power circuits, or per-node domains without splitting into a federation.
spec:
zone: site-a # fallback, still required
zoneFrom:
nodeLabel: topology.kubernetes.io/zone
replication:
factor: 3
zoneRedundancyMode: Maximum # now meaningful — 3 copies in 3 domains
storage:
replicas: 6
Use kubernetes.io/hostname for per-node domains, or a custom label such as example.com/rack for physical racks.
The zone depends on where the pod landed, so it is resolved after scheduling and re-checked on every reconcile. spec.zone is the initial fallback before the Pod is scheduled and whenever the readable Kubernetes Node does not carry the configured label. Once a node has reported an effective status.zone, a transient Pod replacement gap retains that last proven value; this prevents an ordinary rollout from flipping the Garage layout to the fallback zone and back. Failure to read a required Kubernetes Node is different: the resource reports Failed and the operator does not silently mutate the layout using spec.zone. A successfully reconciled node therefore always has a proven zone. The effective value is reported as status.zone on each GarageNode and shown in the ZONE column:
kubectl get garagenodes
NAME CLUSTER ZONE CAPACITY GATEWAY CONNECTED INLAYOUT AGE
garage-storage-0 garage rack-a 500Gi false true true 5m
garage-storage-1 garage rack-b 500Gi false true true 5m
If a pod moves to a Kubernetes Node in a different domain, the layout is updated to match — Garage minimizes the resulting reassignment rather than reshuffling everything. Nodes whose PVCs pin them to a machine will not move at all.
Scope and caveats:
- Operator-managed members. Cluster-level
zoneFromapplies to generated default-group GarageNodes and tonodeLocalPools, including whenstorage.layoutPolicy: Manualdisables only the default group. SetzoneFromdirectly on each user-owned Manual GarageNode. - Storage tier only. It is deliberately not applied to gateway nodes: Garage counts gateway zones toward the
Maximumredundancy target but satisfies that target from storage nodes only, so per-node gateway zones can make every layout apply fail. zoneRedundancyMode: AtLeast(n)requires the label to resolve to at least n distinct values across scheduled storage pods, otherwise layout apply is rejected upstream. The webhook warns about this combination.- Requires the cluster-scoped install (the default). A namespace-scoped install cannot read Nodes, so affected GarageClusters and GarageNodes report
Failedinstead of silently falling back tospec.zone.
Manual Node Layout (GarageNode)
By default, GarageCluster uses layoutPolicy: Auto — the operator generates one operator-owned GarageNode per storage replica (named ), and each GarageNode controller drives its own single-replica StatefulSet. For fine-grained control over individual nodes (per-node zone, capacity, tags, storage class, RPC address, external nodes), set layoutPolicy: Manual and create GarageNode resources directly.
Each GarageNode creates a single-replica StatefulSet and manages that node's layout entry (zone, capacity, tags). Flipping a cluster from Auto → Manual is a one-way hand-off: the operator drops its controllerRef on each GarageNode and the user inherits them. The reverse (Manual → Auto) is rejected by the validating webhook.
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageNode
metadata:
name: storage-node-a
spec:
clusterRef:
name: garage
zone: zone-a
capacity: 500Gi
tags: ["ssd", "high-performance"]
storage:
metadata:
size: 10Gi
data:
size: 500Gi
storageClassName: fast-ssd
Gateway Nodes
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageNode
metadata:
name: gateway-node
spec:
clusterRef:
name: garage
zone: zone-a
gateway: true
storage:
metadata:
size: 1Gi
External Nodes
For nodes running outside Kubernetes (bare-metal, NAS, other clusters):
apiVersion: garage.rajsingh.info/v1beta1
kind: GarageNode
metadata:
name: external-node
spec:
clusterRef:
name: garage
nodeId: "563e1ac825ee3323aa441e72c26d1030d6d4414aeb3dd25287c531e7fc2bc95d"
zone: dc-1
capacity: 1Ti
external:
address: nas.local
port: 3901
External nodes require nodeId (64-hex-char Ed25519 public key). No StatefulSet is created — the operator only manages the layout entry.
Per-Node Overrides
GarageNode supports overriding cluster defaults: image, imageRepository, resources, nodeSelector, tolerations, affinity, podAnnotations, podLabels, priorityClassName, imagePullPolicy, imagePullSecrets, serviceAccountName, securityContext, containerSecurityContext, topologySpreadConstraints, env, envFrom, and logging, plus per-node network.rpcPublicAddr, publicEndpoint, and storage (fsync, snapshots, dataPaths). A node with any of these gets its own ConfigMap instead of the shared cluster config.
Per-Node Maintenance
Set spec.maintenance.suspended: true on a single GarageNode to pause reconciliation of just that node's StatefulSet, ConfigMap, Service, and layout entry — useful for PVC-level work (StorageClass migration, longhorn engine upgrade, disk swap) without the controller fighting you. A Suspended status condition is set while paused, and the finalizer/delete path still runs so a suspended node can be deleted.
Status
kubectl get garagenodes
NAME CLUSTER ZONE CAPACITY GATEWAY CONNECTED INLAYOUT AGE
storage-node-a garage zone-a 500Gi false true true 5m
The controller auto-discovers node IDs from pods, reconciles layout drift (zone/capacity/tags), and handles node removal with replication-safe finalization.
Scaling
GarageCluster supports the Kubernetes scale subresource for the Auto-managed default storage group, enabling kubectl scale and compatibility with autoscalers such as VPA and HPA for that workload.
kubectl scale garagecluster garage --replicas=5
The scale subresource targets .spec.storage.replicas on v1beta2 and .spec.replicas on v1beta1. It controls only the Auto-managed default StatefulSet/PVC group. Node-local pool cardinality follows each storage.nodeLocalPools[].selector; gateway-tier replicas are not exposed through v1beta2 /scale, so adjust spec.gateway.replicas directly. A gateway-only v1beta1 object retains its historical gateway Scale mapping. Dedicated status.scaleReplicas/status.scaleSelector report the actual non-terminating Pods and exact selector for the controllable workload; aggregate status remains separate. Manual storage has no scalable default group—its ordinary GarageNodes are individually owned resources—so HPA/VPA and kubectl scale are unsupported for that shape. A separate fail-closed admission handler runs the full GarageCluster topology validation for /scale, including active drain and prepared scale-down gates.
Because v1beta2 is the preferred discovery version, legacy edge-gateway scaling must name the v1beta1 resource explicitly:
kubectl scale garageclusters.v1beta1.garage.rajsingh.info edge --replicas=5
PVC Retention Policy
By default, PVCs created by a GarageCluster's StatefulSet are not deleted when the cluster is deleted or scaled down. This is intentional: Garage stores your data in those volumes, and automatic deletion would be irreversible.
Storage-member behavior is controlled by spec.storage.pvcRetentionPolicy:
| Field | Value | Behavior |
|-------|-------|----------|
| whenDeleted | Retain (default) | PVCs survive GarageCluster deletion — manual cleanup required |
| whenDeleted | Delete | PVCs are deleted automatically when the GarageCluster is deleted |
| whenScaled | Retain (default) | PVCs for scaled-down pods are kept (allows scaling back up) |
| whenScaled | Delete | PVCs for removed replicas are deleted on scale-down |
For dev/test clusters where you want automatic cleanup:
spec:
storage:
pvcRetentionPolicy:
whenDeleted: Delete
whenScaled: Delete
Requires Kubernetes 1.23+. For production clusters, leave this unset (defaults to Retain) or set whenScaled: Delete only if you're confident scaled-down nodes won't need their data again.
Gateway metadata has a separate spec.gateway.pvcRetentionPolicy. An Auto
unified gateway member owns a single-replica StatefulSet and defaults to
Kubernetes Retain/Retain; an edge gateway keeps its released cluster-level
StatefulSet default of Delete/Delete. An explicit gateway policy applies to
both managed shapes and never changes storage-member claims. A v1beta1 edge
gateway's released spec.storage.pvcRetentionPolicy is losslessly projected to
this field by conversion.
spec:
gateway:
pvcRetentionPolicy:
whenDeleted: Retain
whenScaled: Retain
Choose Delete only when losing that gateway's persisted node_key after its
capacity-less layout role is retired is intentional. Gateway claims contain no
object blocks, but deleting metadata still creates a different Garage identity
on the next start.
Note (Auto mode): automatic storage and gateway topology changes wait until every earlier Garage layout version has leftDraining. The initial storage bootstrap creates the required members together because no Admin API exists yet; every later scale-up or gateway addition admits one per-nodeGarageNodeat a time. Scale-down likewise drains one member at a time and does not begin the next removal until the prior finalizer completes.StorageTopologyReady=FalsereportsAddingMembers,DrainingMember, orWaitingForLayoutSync, and clusterReady=False;StorageScaleDownBlocked=Trueis reserved for a scale-down that would violatereplication.factor. A removed node's single-replica StatefulSet is deleted, so its PVCs are governed bywhenDeleted, notwhenScaled. To reclaim those volumes, setwhenDeleted: Delete.
GarageNode Replacement Cycles
garage.rajsingh.info/cycle=true is a narrow add-before-remove automation for
an established, positive-capacity, StatefulSet-backed GarageNode. Add the
annotation only after the source's exact Pod and 64-hex Garage node ID are
Ready, connected, committed to a settled layout, and observed at the current
generation:
kubectl -n garage-operator-system annotate garagenode garage-storage-a \
garage.rajsingh.info/cycle=true
The operator creates one sibling with a fresh Garage identity and fresh claim
names from the same repeatable PVC templates (or the same explicit EmptyDir
profile), waits for that exact sibling Pod to become Ready and enter the settled
layout, then runs the ordinary cluster-wide drain and block-resync proof before
promoting it. A PVC selector is repeated on the new claim, so static-PV users
must provision distinct replacement PVs in advance.
Automatic cycle deliberately rejects existingClaim, gateways, external
processes, and node-local-pool members. It never infers replacement hardware or
a StorageClass from a bound claim, and it never reuses, clones, snapshots, or
deletes source claims. For SMB, Ceph, manually bound PVCs, exceptional members,
or a different disk profile, explicitly create a second GarageNode with
distinct metadata and data storage, wait for it to synchronize, then drain and
delete the old identity. Node-local membership changes use
spec.storage.nodeLocalPools[].selector and the pool retirement state machine
instead of this annotation.
Progress is durable in status.cyclePhase and the Cycling condition. Removing
the request before the source enters Draining is only a cancellation request:
an already-created sibling must first be explicitly drained and deleted. Once
the source has entered Draining, the annotation and transaction are one-way;
the controller fails closed unless the exact persisted sibling identity appears
in the terminal drain proof.
Multi-HDD Storage
Garage supports striping a node's data across multiple disks. To use it, set spec.storage.data.paths[] instead of spec.storage.data.size — the operator emits one PVC + one volumeMount per path, and renders the matching data_dir TOML array.
spec:
storage:
replicas: 3
metadata:
size: 10Gi
data:
paths:
- path: /data/data0
volume:
size: 1Ti
storageClassName: fast-ssd
- path: /data/data1
volume:
size: 4Ti
storageClassName: bulk-hdd
- path: /mnt/archive
readOnly: true # legacy disk, read-only mount, no capacity
Each storage replica is its own single-replica StatefulSet , so PVCs follow the - convention: data- (e.g. data-0-garage-storage-0-0). The Garage data_dir capacity value is taken from volume.size if set, otherwise from path.capacity, otherwise from the top-level spec.storage.data.size. A readOnly: true path is mounted read-only and emits read_only = true in data_dir — capacity is not required.
Note: Garage usescapacityas a striping weight — blocks are assigned to paths proportionally to each path's capacity. The filesystem enforces the actual size limit, not Garage. In Auto layout mode the cluster spec projects the samepaths[]onto every per-nodeGarageNode, so all storage replicas get identical paths and capacities. For asymmetric per-node disk layouts (e.g. one node with 2×4T, another with 1×8T+1×2T), switch tolayoutPolicy: Manualand setstorage.dataPathsperGarageNode.
Existing clusters: volume source, selector, class/access mode, mount path,
and single↔multi-path topology are immutable whenever a scale transition has
live replicas on either side. The operator supports only in-place size growth
on the same volume. Make topology changes as three separate steps: scale to
zero without changing the template, wait for every exact GarageNode/Pod drain
to finish, then change the template while it remains at zero, and finally
scale up in another unchanged request. Retain PVCs keep their original
selector and class and will be reused; a new selector applies only to a newly
created claim. Intentionally reuse those claims to preserve identity, or,
after the Garage role is fully retired, migrate/remove the exact retained
claims before scaling up. Never delete live metadata/data PVCs: that can lose
node_key or the only local block copies.
selector is supported on default storage metadata/data, each data path,
gateway metadata, and ordinary GarageNode volumes. Releases affected by the
post-#190 projection bug stored cluster selectors but omitted them from newly
generated per-node claims. A non-empty PVC selector matches pre-provisioned PVs;
Kubernetes does not dynamically provision for that claim. Provide one distinct,
access-mode/class-compatible PV for every live metadata/data/path claim, plus
replacement headroom for add-before-remove node cycles. Classless static PVs
usually require an explicit empty storageClassName. This release applies
selectors to new children, cycle replacements, and newly created or recreated
StatefulSets; it does not mutate an existing StatefulSet or reselect an existing
Bound or Pending PVC. Repair an affected Pending claim only through the drained
zero-replica/replacement procedure above. The legacy
volumeClaimTemplateSpec field was never rendered by managed workloads and is
now rejected for new or changed input; an unchanged legacy value is tolerated
with a warning only so it can be removed. Use the explicit PVC fields, or an
ordinary GarageNode with a pre-provisioned existingClaim. To populate new
Auto PVCs from a group-aware snapshot populator, set dataSourceRef on the
volume role (storage.metadata, storage.data, storage.data.paths[].volume,
or gateway.metadata) when creating the cluster.
Custom Container Environment Variables
Both tiers expose env and envFrom for injecting arbitrary env vars into the Garage container, as does GarageNode.spec.env. Built-in vars (GARAGE_NODE_HOST, log sinks) are set first; user entries are appended after, so a user-supplied GARAGE_NODE_HOST would shadow the built-in.
spec:
storage:
env:
- name: GARAGE_ALLOW_WORLD_READABLE_SECRETS
value: "true"
envFrom:
- secretRef:
name: garage-extra-config
gateway:
env:
- name: RUST_BACKTRACE
value: "full"
Reserved variables
These names are operator-reserved and rejected by admission:
| Variable | Also reserved | |---|---| | `GARAGE_CONFI
... (README truncated for length)