8 parts · 8 chapters
Search Engines
Search is the feature users judge most harshly: it must be fast, forgiving of typos, and right about what matters most. Underneath every search engine is the same idea, an inverted index mapping terms to the documents that contain them, plus a scoring function and a lot of careful text processing.
Eight parts: inverted indexes; tokenisation and analysis; scoring with TF-IDF and BM25; Lucene's segments and merges; Elasticsearch and OpenSearch clusters; relevance tuning and evaluation; vector and hybrid search; and a capstone building a small search engine.
indexesPostings lists, term dictionaries, skip lists and compression.
analysisTokenisers, filters, stemming, synonyms, n-grams, local names.
scoringTF-IDF, BM25 and its parameters, boosting.
LuceneImmutable segments, merges, near-real-time refresh.
clustersShards, replicas, mappings, aliases and reindexing.
relevanceJudgements, NDCG, A/B tests, vectors and hybrid ranking.
00
Inverted Indexes
Terms to documents
1 ch · ~8 min01Tokenisation and Analysis
Text becomes terms
1 ch · ~8 min02Scoring: TF-IDF and BM25
Ranking by evidence
1 ch · ~8 min03Lucene Segments and Merges
Immutable segments
1 ch · ~8 min04Elasticsearch and OpenSearch Clusters
Clusters, shards and aliases
1 ch · ~8 min05Relevance Tuning
Measure before tuning
1 ch · ~8 min06Vector and Hybrid Search
Meaning plus exact terms
1 ch · ~8 min07Capstone: A Small Search Engine
The build
1 ch · ~8 minBuilt on DSA and Backend System DesignDSA parts 2, 3 and 8 cover hashing, tries and string algorithms; Backend System Design part 8 designed search and typeahead; AI-native part 7 covered retrieval for LLMs.