Part 1 · 1 chapters · ~8 min

Similarity and Embeddings

Embeddings as learned vectors, cosine similarity and the angle between vectors, dot product and Euclidean distance compared, normalisation, nearest-neighbour search, approximate indexes (HNSW, IVF), storage arithmetic, pgvector, and where embeddings fail (exact identifiers, negation, numbers).

2

Similarity is an angle

code
const cos = (a: number[], b: number[]) => dot(a, b) / (norm(a) * norm(b));
const v = { 'transfer failed': [0.9, 0.0, 0.1, 0.8], 'payment declined': [0.95, 0.0, 0.0, 0.7],
            'jollof rice recipe': [0.0, 0.95, 0.0, 0.0], 'bus fare to Ikeja': [0.3, 0.0, 0.9, 0.0] };
// cos with 'transfer failed': itself 1.000 · payment declined 0.992 · bus fare 0.314 · jollof 0.000

-- Postgres with pgvector: store, index and search
CREATE TABLE docs (id bigserial PRIMARY KEY, body text, embedding vector(1536));
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);
SELECT id, body, 1 - (embedding <=> $1) AS similarity FROM docs ORDER BY embedding <=> $1 LIMIT 5;

Storage arithmetic: a million 1,536-dimension float32 embeddings take 1,000,000 × 1,536 × 4 bytes = 6.14 GB before any index; quantising to int8 cuts that by 4×. Where embeddings fail: exact identifiers (account numbers, error codes), negation ("not refunded") and numbers; combine them with keyword search (hybrid search, Search course P7).

COSINE SIMILARITY, COMPUTED
toy 4-dimension embeddings [money, food, transport, error] vs "transfer failed"
"transfer failed" (itself)1.000"payment declined"0.992"bus fare to Ikeja"0.314"jollof rice recipe"0.000
swipe the figure sideways, or tap expand for full screen
1/4
meaning as position
An embedding model maps text to vectors so that similar meanings land near each other. These toy vectors have four named dimensions; real ones have hundreds to thousands of unnamed ones.
text → vectorsimilar meaning, nearby