Part 0 · 1 chapters · ~8 min
Lexing
The compiler pipeline end to end, tokens and token kinds, the longest-match rule, keywords versus identifiers, skipping whitespace and comments, tracking line and column, lexer errors, lexers as finite automata, and lexer generators versus hand-written lexers.
1
Characters in, tokens out
code
// from modules/byo/repos/compiler/src/lexer.ts
const KEYWORDS = new Set(['fn', 'let', 'if', 'else', 'while', 'return', 'true', 'false', 'int', 'bool', 'print']);
const OPS = ['==', '!=', '<=', '>=', '&&', '||', '+', '-', '*', '/', '%', '<', '>', '!', '='];
if (/[A-Za-z_]/.test(c)) {
let j = i; while (/[A-Za-z0-9_]/.test(src[j] ?? '')) j++;
const text = src.slice(i, j);
out.push({ kind: KEYWORDS.has(text) ? 'kw' : 'ident', text, ...start }); adv(j - i); continue;
}
const op = OPS.find(o => src.startsWith(o, i)); // longest first: '==' before '='
if (op) { out.push({ kind: 'op', text: op, ...start }); adv(op.length); continue; }
throw new CompileError(`unexpected character '${c}'`, line, col);
// run it: cd modules/byo/repos/compiler && npm test → pass 6, fail 0 (Node 25.7)Longest match: the OPS list puts two-character operators first so "==" is never lexed as "=" "=". Keywords are lexed as identifiers and then looked up in a set, the simplest correct approach. Every lexer is a finite automaton (Theory of Computation P1); tools like flex generate them, but most production compilers hand-write theirs for better error messages.
THE BYO COMPILER PIPELINE
471 lines of TypeScript, seven files, six passing tests
swipe the figure sideways, or tap expand for full screen
1/4
lex
Characters become tokens, each carrying its line and column so every later stage can report errors at the right place.
chars → tokenspositions kept