Part 5 · 1 chapters · ~8 min

Relevance Tuning

Relevance as measurable quality, query sets and graded judgements, NDCG, precision and MRR, field boosts, function scores for recency, popularity and distance, typo tolerance with fuzziness, synonyms, learning to rank, click data and its biases, A/B testing search, and zero-result queries.

6

Measure before tuning

code
// a tuned merchant query: exact name, prefix, fuzzy, description, boosted by popularity and distance
GET merchants/_search
{ "query": { "function_score": {
    "query": { "bool": { "should": [
      { "match": { "name": { "query": "yaba pharmcy", "boost": 3 } } },
      { "match": { "name.prefix": { "query": "yaba pharmcy", "boost": 2 } } },
      { "match": { "name": { "query": "yaba pharmcy", "fuzziness": "AUTO" } } },
      { "match": { "description": "yaba pharmcy" } } ] } },
    "functions": [
      { "field_value_factor": { "field": "monthly_txns", "modifier": "log1p", "factor": 0.5 } },
      { "gauss": { "location": { "origin": "6.5158,3.3896", "scale": "3km" } } } ],
    "boost_mode": "sum" } } }
MEASURING RELEVANCE
judgements, offline metrics, then online experiments
query settop + tail queriesjudgementsgraded 0-3 per resultNDCG@10offline scorea ranking changeboost, synonym, analyserA/B testclicks, conversionsship or revert
swipe the figure sideways, or tap expand for full screen
1/5
a query set
Collect representative queries from logs: the most frequent (head) and a sample of rare ones (tail), including queries that returned nothing.
head, tail and zero-result queriesfrom real logs