A hands-on look at how etcd works, why it uses a flat key-space for hierarchies, and how MVCC handles state updates under the hood.
Episode 5 of K8s with Pravesh
Hola Amigos π
Welcome back to K8s-with-Pravesh. In the first few episodes, we covered the holistic journey of Kubernetes internal communicationβexploring what happens when we run kubectl apply, diving deep into the front door of the cluster (the API Server), and clearing up some myths about StatefulSets. Today, we are leveling up our journey and exploring etcd, the storehouse of the Kubernetes cluster.
According to the official docs, βetcd is a consistent and highly-available key value store used as Kubernetes' backing store for all cluster data.β It focuses on being a simple, well-defined, user-facing API (gRPC), secured with automatic TLS and optional client certificate authentication. It is written in Go and uses the Raft consensus algorithm (which allows a cluster of computers to form a single coherent group that can survive individual machine failures) to manage a highly available replicated log. It elects a "leader" node to handle state mutations, which are then replicated over to "follower" nodes. It also follows an odd-number topology because Raft requires a majority quorum to commit data. Thatβs why etcd must run with an odd number of nodes (typically 3 or 5 in production) to handle split-brain scenarios and maintain fault tolerance.
In the context of Kubernetes, it is used to store state files in the form of key-value pairs. This data acts as the single source of truth. Everything Kubernetes knowsβits cluster details, deployments, services, workloads, secretsβis recorded inside etcd. When a user runs kubectl apply -f deployment.yml, the kube-apiserver validates the request and writes the metadata inside etcd. It has a watch mechanism effect where control plane components like the scheduler and controller manager constantly watch etcd using the API server. If a change is detected (like the creation of a new pod or a change in metadata), the new changes, if valid, are committed to the etcd database, and other components will implement that change in the cluster.
Interesting Fact: No component talks directly to etcd; all communication has to go through the API server to ensure security policies, authentication, and strict data validation before data is mutated.
Even though etcd actually uses a key-value databaseβmeaning it doesnβt have real directoriesβit uses a slash-separated naming convention to simulate a deeply nested directory tree. When the kube-apiserver serializes and stores data inside etcd, it uses a default prefix for the root path (/registry). The hierarchy generally splits into two structured paths depending on whether a resource is namespaced or cluster-scoped:
Namespace-Scoped Resources (Most Common): It includes objects related to a specific namespace (Deployments, services, ConfigMaps, secrets, etc.). The pattern is as follows:
/registry/{resource_plural}/{namespace}/{object_name}
- Deployments:
/registry/deployments/default/nginx-demo - Pods:
/registry/pods/default/nginx-demo-6b74467d-9xyz - Secrets:
/registry/secrets/kube-system/default-token-abc12
Cluster-Scoped Resources: For global objects that exist outside any namespace (like nodes, namespaces themselves, cluster roles, and persistent volumes), the namespace segment is dropped:
/registry/{resource-plural}/{object-name}
- Namespaces:
/registry/namespaces/default - Nodes:
/registry/nodes/ip-10-0-1-50.ec2.internal - Persistent Volumes:
/registry/persistentvolumes/pv-data-disk
To see things in action, letβs spin up a Minikube cluster and look at how data is stored inside etcd in the form of slash-separated directories. The commands that we will use in this demonstration are available in the following GitHub repo: K8s_with_Pravesh. Inside the root dir, open part-05-etcd, where you will find the commands in the README.md file.
Once you have your Minikube cluster running, we will create an Nginx deployment using the following command:
kubectl create deployment nginx-demo --image=nginx --replicas=3
Now, to look past the Kubernetes abstraction layer, we will directly exec inside the etcd pod. Since we are using Minikube, we will directly target the control plane pod using its specific certificate paths:
kubectl exec -n kube-system etcd-minikube -- sh -c "ETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/minikube/certs/etcd/ca.crt \
--cert=/var/lib/minikube/certs/etcd/server.crt \
--key=/var/lib/minikube/certs/etcd/server.key \
get /registry/deployments/default/nginx-demo --write-out=fields" | grep -E "Key|Value" | strings
In the output, you will see the key-value pair of our Nginx deployment. Kubernetes uses Protobuf binary to encode data, so some information stays hidden, but you can see messages confirming the successful progress of the Nginx deployment.
Unlike traditional relational databases, Kubernetesβs etcd uses Multi-Version Concurrency Control (MVCC). Every time we modify a Kubernetes object (like updating our Nginx deployment version or image name), etcd doesnβt overwrite itβit increments a global revision number and stores a new version. To see this in a practical demo, use the following command:
kubectl exec -n kube-system etcd-minikube -- sh -c "ETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/minikube/certs/etcd/ca.crt \
--cert=/var/lib/minikube/certs/etcd/server.crt \
--key=/var/lib/minikube/certs/etcd/server.key \
get /registry/deployments/default/nginx-demo --write-out=json" | jq
If we inspect the output, we can see that when we query the specific path (/registry/deployments/default/nginx-demo), the ModRevision is 2814. Now, letβs update our Nginx image version using the following command:
kubectl set image deployment/nginx-demo nginx=nginx:1.25
Now, if we run the same command again and check the ModRevision, we will see something different (3854 in my case). As we update the image version, instead of overwriting, etcd incremented and stored the new version.
Conclusion
That wraps up our deep dive into etcd! Understanding how Kubernetes stores its state behind the scenes, how logical hierarchies are mapped out using flat key prefixes, and how MVCC handles version revisions gives you a whole new perspective when debugging cluster issues. You are no longer just looking at abstract kubectl commandsβyou know exactly what is happening down in the database layer.
I hope you enjoyed this episode of K8s-with-Pravesh. Stay tuned for the next one where we tackle even more core Kubernetes internals!
If you found this helpful, letβs connect:
- YouTube: youtube.com/@pravesh-sudha
- LinkedIn: linkedin.com/in/pravesh-sudha/
- X (Twitter): x.com/praveshstwt
- Medium: medium.com/@programmerpravesh
- Dev.to: dev.to/pravesh_sudha_3c2b0c2b5e0
Happy cluster building, amigos! π


