Part 0 · 1 chapters · ~8 min
Git Internals: Objects, Refs, Packfiles
Content addressing, blobs, trees, commits and annotated tags, inspecting objects with cat-file, refs, HEAD and detached HEAD, the index (staging area), packfiles and delta compression, garbage collection, the SHA-256 transition, and why history is a directed acyclic graph.
1
Look inside a real repository
code
git init -q -b main && printf 'hello\n' > a.txt && git add a.txt && git commit -qm first
git hash-object a.txt # ce013625030ba8dba906f756967f9e9ca394464a
printf 'blob 6\0hello\n' | shasum # ce013625030ba8dba906f756967f9e9ca394464a (the same)
git cat-file -p HEAD # tree 2e81171448eb9f2ee3821e3d447aa6b2fe3ddba1
# author demo <[email protected]> 1791331200 +0000 …
# first
git cat-file -p 'HEAD^{tree}' # 100644 blob ce013625030ba8dba906f756967f9e9ca394464a a.txt
cat .git/HEAD # ref: refs/heads/main
cat .git/refs/heads/main # b8aa7c0db9bc29b8e2bccc93e87d8ff42f3cd0a2
find .git/objects -type f | wc -l # 3: one blob, one tree, one commitThe index (.git/index) is the proposed next tree: git add copies content into blobs and records them there; git commit turns the index into a tree. Packfiles: loose objects are later packed into .git/objects/pack with delta compression between similar objects (git gc), which is why clones are much smaller than the sum of all versions. Git is moving to SHA-256 object names (git init --object-format=sha256) because SHA-1 collisions became practical in 2017.
GIT'S OBJECT MODEL, FROM A REAL REPO
git 2.46, one file, one commit
swipe the figure sideways, or tap expand for full screen
1/4
blobs
File contents are stored as blobs named by the SHA-1 of "blob <size>\0<content>". hello\n became ce013625…; shasum over the same bytes printed the same hash.
content → hash → nameidentical content stored once