Part 0 · 1 chapters · ~8 min

Tokenisation

Why models need tokens, byte-pair encoding and how vocabularies are learned, byte-level fallback, measured token counts for English, Nigerian names, the naira sign and Yoruba, how digits split, token counts and cost, context windows, and special tokens and chat templates.

1

Text to integers

code
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
for s in ["transfer failed", "Send ₦5,000 to Adéọlá", "Ẹ káàárọ̀, ṣé dáadáa ni?", "1234567890"]:
    ids = enc.encode(s); print(len(s), len(ids), [enc.decode([i]) for i in ids])
# 15  2  ['transfer', ' failed']
# 21 11  ['Send', ' �', '�', '5', ',', '000', ' to', ' Ad', 'é', 'ọ', 'lá']      ← ₦ is two byte tokens
# 24 15  ['Ẹ', ' ká', 'à', 'ár', 'ọ', '̀', ',', ' ṣ', 'é', ' dá', 'ad', 'á', 'a', ' ni', '?']
# 10  4  ['123', '456', '789', '0']

Byte-pair encoding starts from bytes and repeatedly merges the most frequent adjacent pair in the training corpus into a new token, until the vocabulary reaches its target size (about 200,000 for o200k_base). Frequent strings become single tokens; anything else falls back to smaller pieces, down to raw bytes, so every input is representable. Chat templates wrap messages in special tokens marking roles (system, user, assistant); the model only ever sees one long token sequence.

TOKENS PER STRING, MEASURED
OpenAI o200k_base tokenizer via tiktoken
"transfer failed" (15 chars)2 tokens"Good morning, how are you?" (26)7 tokens"Send 5000 naira to Adeola" (25)9 tokens"Send ₦5,000 to Adéọlá" (21)11 tokens"Ẹ káàárọ̀, ṣé dáadáa ni?" (24)15 tokens
swipe the figure sideways, or tap expand for full screen
1/4
common words
Frequent English words are single tokens: "transfer failed" is 2 tokens, "Good morning, how are you?" 7.
common words = 1 tokenEnglish is cheap