Part 6 · 1 chapters · ~8 min

Evaluation

Why public benchmarks do not predict your feature's quality, contamination, building golden sets from production data, deterministic checks, LLM-as-judge with rubrics and calibration, pairwise comparisons, retrieval evaluation (recall@k, groundedness), online metrics, regression gating in CI, and evaluating cost and latency alongside quality.

7

An eval harness

code
// evals/ticket-classify.eval.ts: runs in CI on every prompt or model change
const cases = loadJsonl('evals/golden/tickets.jsonl');            // 300 real tickets, labelled
let pass = 0, cost = 0; const lat: number[] = [];
for (const c of cases) {
  const t0 = Date.now(); const r = await classify(c.input); lat.push(Date.now() - t0); cost += r.usage.costUsd;
  const ok = r.output.category === c.expected.category && schema.safeParse(r.output).success;
  if (ok) pass++; else report(c, r.output);
}
const acc = pass / cases.length;
console.log({ acc, p95ms: pct(lat, 95), costPer1k: cost / cases.length * 1000 });
if (acc < 0.92) process.exit(1);                                 // regression gate

Retrieval features need two layers of evaluation: did retrieval find the right documents (recall@k against labelled relevant chunks), and did the answer use them faithfully (groundedness: every claim supported by a retrieved passage)? Track quality, latency and cost together; a change that gains 1 point of accuracy at 3× cost may not be worth shipping.

EVALUATING LLM FEATURES
build your own evals before you change anything
golden set50-500 real inputs with expectedoutputs or rubrics, fromproduction.exact checksJSON validity, schema, requiredfields, regex, numeric tolerance.model gradersAn LLM judges against a rubric;calibrate it against human labels.pairwiseCompare old vs new outputs side byside; less noisy than scores.onlineUser feedback, edits, escalations,task completion in production.regression gateRun evals in CI on every prompt ormodel change.
swipe the figure sideways, or tap expand for full screen
1/4
golden sets
Public benchmarks measure general ability; your feature needs its own set from real traffic, including the hard and embarrassing cases.
real inputs, expected outputsyour own, not leaderboards