Tokenisation
Why models need tokens, byte-pair encoding and how vocabularies are learned, byte-level fallback, measured token counts for English, Nigerian names, the naira sign and Yoruba, how digits split, token counts and cost, context windows, and special tokens and chat templates.
Text to integers
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
for s in ["transfer failed", "Send ₦5,000 to Adéọlá", "Ẹ káàárọ̀, ṣé dáadáa ni?", "1234567890"]:
ids = enc.encode(s); print(len(s), len(ids), [enc.decode([i]) for i in ids])
# 15 2 ['transfer', ' failed']
# 21 11 ['Send', ' �', '�', '5', ',', '000', ' to', ' Ad', 'é', 'ọ', 'lá'] ← ₦ is two byte tokens
# 24 15 ['Ẹ', ' ká', 'à', 'ár', 'ọ', '̀', ',', ' ṣ', 'é', ' dá', 'ad', 'á', 'a', ' ni', '?']
# 10 4 ['123', '456', '789', '0']Byte-pair encoding starts from bytes and repeatedly merges the most frequent adjacent pair in the training corpus into a new token, until the vocabulary reaches its target size (about 200,000 for o200k_base). Frequent strings become single tokens; anything else falls back to smaller pieces, down to raw bytes, so every input is representable. Chat templates wrap messages in special tokens marking roles (system, user, assistant); the model only ever sees one long token sequence.