Profile
Back to NewsBack
Dev.to 15 min
Reader Mode
Git Isn't a Diff Tracker: How Blobs, Trees, DAG Commits, and the Index Actually Work Under the Hood

Git Isn't a Diff Tracker: How Blobs, Trees, DAG Commits, and the Index Actually Work Under the Hood

17 hours ago

Ask most developers how Git works, and you will hear a familiar explanation:

"Git tracks changes. When you edit a file and commit, Git calculates the diff between the old file and the new file and saves the patch."

It is a completely reasonable mental model. After all, when you run git diff or git show, the terminal prints green additions and red deletions. Older version control systems like CVS and Subversion (SVN) worked exactly that way: they stored a base file and appended delta patches across time.

There is only one problem: Git does not work that way at all.

Git does not track file diffs. Git does not store patches. Git does not even have an internal concept of "file renames".

Underneath its porcelain command line interface, Git is something radically simpler and far more elegant: an immutable, content-addressed key-value object store sitting underneath a Directed Acyclic Graph (DAG) of snapshot trees.

When developers treat Git as a magical incantation machine, commands like git rebase, git reset --mixed, or a detached HEAD state feel like minefields. But the moment you look inside the .git folder and understand how Git represents data, the mystery dissolves.

Here is what actually happens under the hood when you use Git.


1. Peeking Inside .git: The Content-Addressed Object Store

When you type git init in an empty folder, Git does not set up a server or connect to a daemon. It creates a single hidden directory named .git.

Inside that directory, the most important folder is .git/objects.

This folder is Git's database. Every piece of data in your project, every file, every directory listing, every commit message, and every tag, is stored here as one of four fundamental object types:

  1. Blob: Stores raw file contents (uncompressed byte streams). It does not store file names, paths, or file permissions.
  2. Tree: Represents a directory. It maps file names and file permissions to the cryptographic hashes of blobs and child trees.
  3. Commit: A snapshot pointer. It references a top-level root tree, zero or more parent commit hashes, an author, a committer, and a commit message.
  4. Annotated Tag: A permanent pointer to a specific commit, containing a tagger name, date, and verification message.

The Four Git Object Types in the Object Database

The Cryptographic Hash (Content-Addressing)

Git identifies every object by its cryptographic hash. In standard Git, this is a 40-character hexadecimal SHA-1 string (newer versions can also use SHA-256).

Let us demonstrate this with raw plumbing commands.

If you create a file with the content hello world\n and pass it to Git's low-level hash calculation tool:

$ echo "hello world" | git hash-object -w --stdin
3b18e512dba79e4c8300dd08aeb37f8e728b8dad

The -w flag instructs Git to write the object directly into .git/objects.

How did Git compute 3b18e512dba79e4c8300dd08aeb37f8e728b8dad?

Git prepends a standardized header to the raw data: the object type (blob), a space, the content length in bytes, and a null byte (\0). Then it passes the concatenated buffer through the SHA algorithm:

header  = "blob 12\0"
content = "hello world\n"
store   = header + content

sha1("blob 12\0hello world\n") = 3b18e512dba79e4c8300dd08aeb37f8e728b8dad

Notice what is missing from this calculation: the file name.

If you name that file README.md, main.py, or config.json, the SHA hash is identical. If five different directories in your project contain an identical file, Git stores that blob exactly once.

How Git Organizes the File System

Inspect .git/objects after writing that blob, and you will see:

.git/objects/
├── 3b/
│   └── 18e512dba79e4c8300dd08aeb37f8e728b8dad

Git takes the first two characters of the 40-character hash (3b) and creates a subfolder. The remaining 38 characters become the filename. This split prevents operating systems from choking when thousands of objects accumulate in a single directory.

The file on disk is compressed using zlib. You can inspect its contents and type at any time using Git's inspection tool, git cat-file:

# Check the object type
$ git cat-file -t 3b18e512
blob

# Print the uncompressed contents
$ git cat-file -p 3b18e512
hello world

2. The Snapshot Architecture: Why Git Never Stores Diffs

Because older version control tools stored delta diffs, many developers assume Git calculates differences when saving new commits.

In reality, Git stores complete snapshots.

When you commit a change, Git creates a new Tree object representing your root folder. That Tree object contains a complete manifest of every file and folder in the project.

Delta-Based VCS vs Git Snapshot Trees

The Efficiency of Immutable Trees

You might wonder: If Git stores full snapshots rather than diffs, why doesn't my repository size explode into gigabytes after twenty commits?

The secret is content-addressable structural sharing.

Suppose your repository contains 1,000 files across 20 folders. You open docs/guide.md, fix a single typo, and run git commit.

Here is what Git does:

  1. It hashes the new docs/guide.md and writes one new blob into .git/objects.
  2. It creates a new Tree object for the docs/ folder. This tree points to the new blob for guide.md, but reuses the existing hashes for all unchanged files in docs/.
  3. It creates a new root Tree object. This root tree points to the new docs/ tree, but reuses the exact same hashes for the other 19 unchanged folder trees.
  4. It creates a new Commit object pointing to the new root tree.

999 files were not copied or re-hashed. Their pointers remained identical across the tree hierarchy.

Because blobs and trees are immutable, Git reuses them indefinitely. If a 100MB asset never changes across 500 commits, it takes up 100MB of disk space, not 50GB.

When Does Git Actually Use Deltas? (Packfiles)

Git does use delta compression, but only during optimization, never during normal commit creation.

As you work, loose objects accumulate in .git/objects. Periodically, or when you run git gc (garbage collection) or git push, Git packs loose objects into a Packfile (.pack) alongside an index file (.idx).

Inside a packfile, Git runs an intelligent sliding-window delta compression algorithm. It compares files with similar sizes and paths, identifies common patterns, and stores one file as a full version and the others as reverse deltas.

This two-tier architecture gives Git the best of both worlds: O(1) instantaneous snapshot creation during development, and minimal network transfer during pushes and pulls.


3. The Three Trees: Working Directory, Index, and HEAD

To understand how commands like add, commit, checkout, and reset actually work, you must visualize Git as managing three separate trees.

The Three Trees: Working Directory, Index, and HEAD

1. The Working Directory (Filesystem)

This is the physical sandbox on your hard drive where you edit files in your editor. These are regular, uncompressed files that your operating system and tools interact with directly.

2. The Index (The Staging Area)

The index is not an abstract concept; it is a physical binary file located at .git/index.

The index is a cache that pre-computes what your next commit will look like. When you execute git add path/to/file:

  • Git reads the file from your working directory.
  • It writes a new compressed blob into .git/objects.
  • It records the file path, permissions, and the new blob's SHA hash inside .git/index.

The staging area is effectively an in-memory draft of a Tree object, waiting to be sealed into a commit.

3. The HEAD (Last Committed State)

HEAD is a reference pointer. It tracks your current location in the repository's commit graph. Most of the time, HEAD points to a branch name, which in turn points to the most recent commit object in that branch.

Demystifying git reset: The Three-Tree Slider

Once you understand the three trees, Git's most feared command, git reset, becomes completely intuitive.

git reset <commit-hash> is simply a slider that allows you to choose how many of the three trees to roll backward:

  • git reset --soft HEAD~1: Moves only HEAD backward by one commit. Your Index and Working Directory remain untouched. All changes from that commit remain staged, ready to be recommitted.
  • git reset --mixed HEAD~1 (the default): Moves HEAD and updates the Index to match that previous commit. Your Working Directory remains untouched. Your changes are safe on disk, but appear as unstaged modifications.
  • git reset --hard HEAD~1: Moves HEAD, overwrites the Index, and overwrites your Working Directory. Any uncommitted changes on disk are permanently discarded.

Instead of memorizing confusing command-line flags, visualize moving a pointer across the three tiers.


4. Branches and Merging: Pointers in a Directed Acyclic Graph

In many traditional version control systems, creating a branch was an expensive operation: it copied the entire project directory tree into a separate /branches/feature folder.

In Git, a branch is not a copy of your files. A branch is just a 41-byte text file.

The Anatomy of a Branch Reference

Look inside .git/refs/heads/:

$ ls .git/refs/heads/
main
feature-login

Open .git/refs/heads/main in a text editor:

7c841a0293dbfc9d6138fbd8f11075677e4e892c

It is a single 40-character hex string followed by a newline character.

When you create a new branch with git branch feature-auth, Git does not duplicate a single line of source code. It creates a new 41-byte text file named feature-auth containing the exact commit SHA that your current branch points to.

This is why branch creation in Git takes less than one millisecond, regardless of whether your project contains ten files or ten million files.

Branching, Fast-Forward Merges, and 3-Way Graph Merges

What is HEAD?

If branches are just pointers to commit objects, how does Git know which branch you are currently working on?

Inspect .git/HEAD:

$ cat .git/HEAD
ref: refs/heads/main

HEAD is a symbolic pointer that references your current branch file.

When you run git commit:

  1. Git writes the new tree and commit object.
  2. The new commit points to the old commit as its parent.
  3. Git updates .git/refs/heads/main to point to the new commit's hash.
  4. HEAD does not move; it still points to main, which now points to the new commit.

What is a "Detached HEAD"?

When you run git checkout <commit-hash> instead of git checkout <branch-name>, Git writes the raw commit hash directly into .git/HEAD:

$ cat .git/HEAD
7c841a0293dbfc9d6138fbd8f11075677e4e892c

You are now in a Detached HEAD state.

It sounds intimidating, but it simply means: HEAD is pointing directly to a commit object rather than to a branch pointer.

If you make commits in this state, the new commits are created normally, but no branch reference advances with them. If you switch branches, those commits become unreachable by name, which leads us to merging and rebasing.

Merges: Fast-Forward vs Three-Way Merges

When you merge branch B into branch A, Git evaluates the commit graph:

  • Fast-Forward Merge: If branch A has not received any new commits since branch B branched off, Git does not create a merge commit. It simply moves the pointer in .git/refs/heads/A forward to match branch B.
  • Three-Way Merge: If both branches have diverged with independent commits, Git must perform a 3-way merge. It identifies three distinct snapshots:
    1. The tip of Branch A (Ours).
    2. The tip of Branch B (Theirs).
    3. The Best Common Ancestor (the point where the two branches diverged).

By comparing both tips against the common ancestor, Git determines which changes were made on which branch. It constructs a new Tree combining both sets of edits and writes a Merge Commit containing two parent hashes.


5. Rebase, Cherry-Pick, and the Safety Net of the Reflog

Few topics generate as much confusion as the difference between git merge and git rebase.

Both commands integrate changes from one branch into another, but they do it in fundamentally different ways in the commit graph.

Rebase vs Merge and the Reflog Safety Net

How git rebase Actually Works

When you run:

$ git checkout feature
$ git rebase main

Git does not modify your existing commits. Remember: Git objects are cryptographically immutable. If you change a commit's parent, its SHA hash changes completely.

Instead, Git performs the following sequence:

  1. It identifies all commits unique to feature since it diverged from main.
  2. It temporarily saves the diffs of those commits into memory.
  3. It resets the feature branch pointer to point directly to the current tip of main.
  4. It replays each saved change onto the new base one by one, generating brand-new commit objects with brand-new SHA hashes.
  5. The original commits are left behind in the object database as orphaned nodes.

Rebasing creates a clean, linear project history, but it rewrites commit history. This is why the universal golden rule of Git exists: never rebase commits that have been pushed to a shared public branch.

The Ultimate Safety Net: The Reflog

Suppose you accidentally ran a disastrous command:

$ git reset --hard HEAD~5

Five valuable commits appear to have vanished. Your working directory is reverted. git log shows no trace of your work.

Do not panic. Git almost never deletes committed data immediately.

Every time HEAD or a branch pointer moves in your local repository, Git logs the previous position in .git/logs/HEAD. This log is called the Reflog (Reference Log).

Run:

$ git reflog

You will see a chronological record of every pointer movement:

7c841a0 (HEAD -> main) HEAD@{0}: reset: moving to HEAD~5
9e4b102 HEAD@{1}: commit: Implement user authentication
2a8f331 HEAD@{2}: commit: Add database connection pool

The commits were never destroyed; only the pointer was moved.

To recover your "lost" commits instantly, simply point a branch back to the commit hash recorded in the reflog:

$ git branch recovery-branch 9e4b102

Your five commits, along with all their trees and blobs, are restored immediately.

Git keeps unreachable objects in .git/objects for at least 30 to 90 days before running automatic garbage collection (git prune). As long as you committed your work at least once, it is virtually impossible to permanently lose it.


6. Three Common Git Traps (And the Internal Fixes)

Understanding Git's internal architecture immediately provides solutions to the three most frustrating problems developers encounter:

Trap 1: Accidentally Committing a Large File or Secret

A developer accidentally commits a 200MB database dump or an API key file, realizes the mistake, runs git rm data.sql, and commits again.

The repository size remains 200MB, and the file is still visible in history.

Why it happens: Running git rm only removes the file from future Tree objects. The previous Commit object still references the original Tree, which still points to the 200MB blob in .git/objects.

The fix: To completely remove an object from history, you must rewrite the commit DAG using tools like git-filter-repo or BFG Repo-Cleaner, and then force garbage collection:

# Modern replacement for git filter-branch
git-filter-repo --invert-paths --path data.sql
git reflog expire --expire=now --all
git gc --prune=now --aggressive

Trap 2: Merge Conflict Panic

Developers often panic during merge conflicts because they look only at the conflict markers:

<<<<<<< HEAD
const API_URL = "https://api.v1.prod.com";
=======
const API_URL = "https://api.v2.internal.net";
>>>>>>> feature-branch

Why it happens: A conflict occurs because both branches modified the exact same lines of code relative to their common ancestor snapshot.

The internal fix: Use Git's 3-way merge diff style to inspect what the ancestor originally had:

git config --global merge.conflictstyle zdiff3

Now conflict markers display the base version between the two branches:

<<<<<<< HEAD
const API_URL = "https://api.v1.prod.com";
||||||| base
const API_URL = "http://localhost:3000";
=======
const API_URL = "https://api.v2.internal.net";
>>>>>>> feature-branch

Seeing the common ancestor makes resolving the conflict straightforward: you immediately see what each side was trying to achieve.

Trap 3: Blind Force Pushing (git push --force)

When a developer rebases locally, pushing to remote fails because the remote branch's history has diverged. The common instinct is to run git push --force.

Why it happens: --force overwrites the remote ref pointer unconditionally, potentially obliterating commits that a teammate pushed while you were rebasing.

The internal fix: Always use --force-with-lease:

git push --force-with-lease

This flag checks the remote ref before overwriting. If another developer pushed commits to that branch in the meantime, the push will abort, protecting their work while letting you rebase cleanly.


The Mental Shift

The difference between struggling with Git and mastering it comes down to a shift in perception:

Stop thinking of Git as an opaque command-line tool that saves edits and file diffs.

Git is a simple, elegant content-addressed database.

  • Files are Blobs, addressed by their cryptographic hash.
  • Folders are Trees, mapping names to hashes.
  • Snapshots are Commits, pointing to trees and parents in a Directed Acyclic Graph.
  • Branches and tags are just lightweight 41-byte text pointers.
  • Merges and rebases are operations that navigate and reshape the graph.

Once you visualize the underlying graph, confusing commands become predictable, broken branches are easily untangled, and the fear of losing code disappears completely.

How did understanding Git internals change the way you navigate daily workflows? Have you ever had to rescue a production branch using the reflog, or untangle a complex rebase conflict using plumbing commands? Share your stories and favorite diagnostic commands in the comments below.

Chat with me