Part 1 · 1 chapters · ~8 min
Embeddings
The embedding table as a learned lookup, embedding dimension and parameter count, geometry and similarity, positional information (learned, sinusoidal, rotary), the unembedding layer that turns vectors back into token scores, weight tying, and the difference between token embeddings and text-embedding models.
2
A lookup, then geometry
code
# an embedding layer is a matrix indexed by token id import torch vocab, d = 128_256, 4_096 # Llama 3.1 8B sizes E = torch.nn.Embedding(vocab, d) # 128,256 × 4,096 ≈ 525 million parameters x = E(torch.tensor([[791, 1404, 3477]])) # shape (1, 3, 4096): one vector per token # at the other end, the unembedding (often the same matrix, "tied") gives a score per vocabulary entry logits = h_final @ E.weight.T # (1, 3, 128256) → softmax → next-token probabilities
The vocabulary-sized matrices at both ends are a large share of a small model's parameters; this is one reason vocabularies are not made arbitrarily large.
FROM TOKEN IDS TO VECTORS
the first layer of every transformer
swipe the figure sideways, or tap expand for full screen
1/4
lookup
Each token id selects a row of a learned table: a vector of d numbers (4,096 in Llama 3.1 8B). Nothing more than an array index.
id → row of a matrixlearned during training