Part 8 · 2 chapters · ~12 min

Search and Typeahead

Full-text search over business data with an inverted index fed by CDC, relevance with BM25 plus business signals, typeahead under 50 ms per keystroke, index freshness, and reindexing without downtime.

17

Brief, questions and numbers

the brief
  1. Users search merchants, transactions and help articles, with suggestions as they type.
questionanswer we assume
corpus?50M merchants and products, 2B transactions (per-user scope)
latency?search p99 300 ms, typeahead p99 50 ms
freshness?seconds
languages?English, with Nigerian names and pidgin terms
filters?category, location, date ranges
code
queries    = 5,000/s search, 30,000/s typeahead keystrokes
index size = 50M docs × 2 KB ≈ 100 GB primary + replicas
per-user transaction search: filter by user first; index only recent months
SEARCH AND TYPEAHEAD
an inverted index fed by change events, and a prefix service for suggestions
source of truthPostgresCDC / outboxchange eventsindexerbuild documentssearch clusterOpenSearch / Elasticsearchtypeaheadprefix trie / edge n-gramssearch APIquery, rank, filter
swipe the figure sideways, or tap expand for full screen
1/5
feed the index
Changes flow from the database through CDC or an outbox to an indexer that builds search documents (denormalised: merchant name, category, location).
CDC feeds the indexerdocuments are denormalised
18

v1, the break, and v2

v1. SQL LIKE '%term%' queries on the merchants table, run on every keystroke.

The break. Leading-wildcard LIKE cannot use B-tree indexes: full scans per keystroke, no relevance ranking, no typo tolerance.

v2. An OpenSearch cluster fed by CDC with denormalised documents, analysers for names and local terms, a typeahead service with precomputed popular prefixes, per-user transaction search scoped by filters, and zero-downtime reindexing through index aliases.

the sentence
v2 buys relevant, fast, typo-tolerant search, and pays with a second copy of the data that lags the source and a cluster to operate.