Part 2 · 1 chapters · ~8 min
Scoring: TF-IDF and BM25
Relevance as a ranking problem, term frequency and inverse document frequency, BM25 with saturation (k1) and length normalisation (b) worked numerically, field boosts and multi-field queries, function scores for recency and popularity, and explaining a score.
3
Ranking by evidence
code
BM25 per term: idf(t) × tf × (k1 + 1) / (tf + k1 × (1 − b + b × docLen / avgDocLen))
idf(t) = ln(1 + (N − n_t + 0.5) / (n_t + 0.5)) N documents, n_t containing t
with k1 = 1.2, b = 0.75 and an average-length document: tf 1 → 1.00, 2 → 1.38, 5 → 1.77, 10 → 1.96, 50 → 2.15 (× idf)
GET merchants/_search { "explain": true, "query": { "multi_match": { "query": "yaba pharmacy",
"fields": ["name^3", "description"], "type": "best_fields" } } } // name matches weigh 3×BM25: TERM FREQUENCY SATURATES
score contribution of one term as its count in a document grows (k1 = 1.2, average-length document)
swipe the figure sideways, or tap expand for full screen
1/4
TF-IDF
Classic scoring multiplies term frequency (how often the term appears in the document) by inverse document frequency (how rare the term is across all documents). Rare terms matter more; "pharmacy" outweighs "the".
frequency × rarityrare terms count more