Git Internals

Most Git confusion (detached HEAD, "lost" commits, rebase rewriting hashes, why a branch is "just a pointer") disappears once you see the data model underneath. At its core, Git is a content-addressed object database: every file version, directory listing, and commit is stored as an object named by the hash of its contents. On top of that sit refs, human-readable names like main that point at commits, and the index, a staging area describing the next commit.

That's almost the whole design. Branches, tags, merges, rebases, and stashes are operations on those few structures. This page walks through them with the plumbing commands that let you inspect them directly.

TL;DR

Quick Example

Peek inside a repository with plumbing commands:

Core Concepts

The Object Database

Everything lives in .git/objects. Each object is stored under its hash:

Because objects are content-addressed, identical files anywhere in history share one blob, renames cost nothing (the tree changes, not the blob), and an object can never be modified. A different content means a different hash, and so a different object.

Commits Are Snapshots, Chained by Hashes

A commit records the complete tree of the project at that moment, not a diff. Diffs are computed on demand by comparing trees. Each commit includes its parent's hash, so each commit's ID depends on its entire ancestry, like a hash-chained linked list. This is why:

Refs, Branches, and HEAD

A ref is a name for a commit hash, stored as a file under .git/refs/ or packed into .git/packed-refs:

Committing on a branch creates a new commit and moves the branch ref to it. That's all "being on a branch" means. HEAD normally contains ref: refs/heads/main (a symbolic ref). When you check out a commit directly, HEAD contains a raw hash: detached HEAD. New commits made there aren't on any branch, and they're easy to lose unless you create one (git switch -c rescue).

The reflog records every position each ref and HEAD have held. It's the safety net for recovering "lost" commits. See undoing things in Git.

The Index (Staging Area)

.git/index is a binary file listing every tracked path with its blob hash, mode, and file-system metadata. git add writes the file's blob into the object database and updates the index entry; git commit turns the index into tree objects and creates a commit pointing at the root tree. Three trees are always in play:

Understanding these three trees explains what reset --soft/--mixed/--hard and restore --staged do.

Packfiles and Garbage Collection

New objects start as individually zlib-compressed loose objects. Periodically (git gc, triggered automatically), Git packs them into packfiles with delta compression between similar objects, which is why repositories with long histories stay compact. Network transfers also send packfiles. Objects unreachable from any ref or reflog entry are pruned after a grace period (two weeks by default), so recently "lost" commits remain recoverable.

Plumbing vs Porcelain

Git's user-facing commands (commit, switch, merge) are porcelain. They're built on plumbing commands meant for scripts and for understanding:

Use plumbing in scripts: its output format is stable, while porcelain output can change between versions.

Best Practices

Think in Pointers

When an operation feels risky, ask which refs it moves and which objects it creates. reset moves a branch ref, commit creates objects and moves a ref, and rebase creates new commits and moves the ref to them. Old objects stay reachable through the reflog for a while.

Prefer Annotated, Signed Tags for Releases

Annotated tags are real objects with a tagger, date, and message, and can be GPG or SSH signed. Lightweight tags are just refs. Use annotated tags for releases (git tag -a v2.0.0 -m "…").

Keep Large Binaries Out of History

Every version of every file is stored forever and cloned by everyone. Large, frequently changing binaries bloat repositories permanently. Use Git LFS or external artifact storage for them.

Common Mistakes

Committing on a Detached HEAD and Switching Away

Before leaving, run git switch -c my-work to put a branch ref on them. If you already left, find them in git reflog.

Believing git rm Removes a File From History

Deleting a file in a new commit removes it from future snapshots; every old commit still references the blob. Purging a leaked secret requires rewriting history (git filter-repo) and force-pushing, and you must still rotate the secret, because clones and forks may already have it.

Manually Editing .git Internals

Hand-editing files in .git/ can corrupt a repository in subtle ways. Use plumbing commands (update-ref, symbolic-ref), which validate and log their changes.

FAQ

Does Git store diffs or full snapshots?

Conceptually, full snapshots: each commit points to a complete tree. Physically, packfiles store many objects as deltas against similar objects for compression. You get the simplicity of snapshots with the storage efficiency of diffs.

Is Git's use of SHA-1 a security problem?

Git uses a hardened SHA-1 variant that detects known collision attacks, and newer versions support SHA-256 object formats for new repositories. For supply-chain integrity, sign commits and tags and verify signatures, rather than relying on hashes alone.

Why is a branch so cheap to create?

Because a branch is just a ref: a tiny file (or a line in packed-refs) containing a 40-character hash. Creating one copies nothing; it names an existing commit.

What exactly does git gc delete?

It packs loose objects and removes objects that are unreachable from every ref, index entry, and reflog entry, and older than the prune grace period. Reachable history is never deleted, and recently abandoned commits survive until their reflog entries expire.

Related Topics

References