Part 6 · 1 chapters · ~8 min
Evaluation
Why public benchmarks do not predict your feature's quality, contamination, building golden sets from production data, deterministic checks, LLM-as-judge with rubrics and calibration, pairwise comparisons, retrieval evaluation (recall@k, groundedness), online metrics, regression gating in CI, and evaluating cost and latency alongside quality.
7
An eval harness
code
// evals/ticket-classify.eval.ts: runs in CI on every prompt or model change
const cases = loadJsonl('evals/golden/tickets.jsonl'); // 300 real tickets, labelled
let pass = 0, cost = 0; const lat: number[] = [];
for (const c of cases) {
const t0 = Date.now(); const r = await classify(c.input); lat.push(Date.now() - t0); cost += r.usage.costUsd;
const ok = r.output.category === c.expected.category && schema.safeParse(r.output).success;
if (ok) pass++; else report(c, r.output);
}
const acc = pass / cases.length;
console.log({ acc, p95ms: pct(lat, 95), costPer1k: cost / cases.length * 1000 });
if (acc < 0.92) process.exit(1); // regression gateRetrieval features need two layers of evaluation: did retrieval find the right documents (recall@k against labelled relevant chunks), and did the answer use them faithfully (groundedness: every claim supported by a retrieved passage)? Track quality, latency and cost together; a change that gains 1 point of accuracy at 3× cost may not be worth shipping.
EVALUATING LLM FEATURES
build your own evals before you change anything
swipe the figure sideways, or tap expand for full screen
1/4
golden sets
Public benchmarks measure general ability; your feature needs its own set from real traffic, including the hard and embarrassing cases.
real inputs, expected outputsyour own, not leaderboards