Part 1 · 1 chapters · ~8 min
Tokenisation and Analysis
The analysis pipeline (character filters, tokenisers, token filters), lowercasing and ASCII folding for names with diacritics, stemming and lemmatisation, stop words, synonyms at index and query time, n-grams and edge n-grams for partial matching, language analysers, and per-field analysis.
2
Text becomes terms
code
// Elasticsearch / OpenSearch: a names analyser and a prose analyser
PUT merchants
{ "settings": { "analysis": {
"filter": { "edge": { "type": "edge_ngram", "min_gram": 2, "max_gram": 15 } },
"analyzer": {
"names": { "tokenizer": "standard", "filter": ["lowercase", "asciifolding"] },
"prefix": { "tokenizer": "standard", "filter": ["lowercase", "asciifolding", "edge"] } } } },
"mappings": { "properties": {
"name": { "type": "text", "analyzer": "names", "fields": { "prefix": { "type": "text", "analyzer": "prefix", "search_analyzer": "names" } } },
"description": { "type": "text", "analyzer": "english" },
"category": { "type": "keyword" } } } }
POST merchants/_analyze { "analyzer": "names", "text": "Adénìjí's Phármacy" } // [adeniji's, pharmacy]ANALYSIS: TEXT TO TERMS
the same pipeline runs at index time and query time
swipe the figure sideways, or tap expand for full screen
1/5
one pipeline, twice
Analysis turns text into terms. The same analyser must run at index time and query time, or a query term can never match an indexed term.
same analyser for index and querymismatch = no matches