Part 8 · 1 chapters · ~12 min

Evals in Depth

An eval as a test suite for non-deterministic behaviour: cases written as expected properties, code graders before model rubrics, a threshold with must-pass cases, version-to-version diffs, growing the set from production failures with a held-out slice, and running it in CI, nightly and in mock mode.

9

Cases, graders, gates and diffs

code
// an eval harness in miniature
type Case = { id: string; input: string; tags: string[]; mustPass?: boolean;
              expect: { cites?: string[]; contains?: string[]; absent?: string[]; refuses?: boolean } };

async function grade(c: Case, out: Answer): Promise<{ pass: boolean; why: string[] }> {
  const why: string[] = [];
  for (const s of c.expect.cites ?? []) if (!out.citations.includes(s)) why.push(`missing citation ${s}`);
  for (const s of c.expect.contains ?? []) if (!out.text.includes(s)) why.push(`missing "${s}"`);
  for (const s of c.expect.absent ?? []) if (out.text.includes(s)) why.push(`forbidden "${s}"`);
  if (c.expect.refuses !== undefined && out.refused !== c.expect.refuses) why.push('refusal mismatch');
  return { pass: why.length === 0, why };
}

async function runEval(cases: Case[], system: System, runs = 3) {
  const results = [];
  for (const c of cases) {
    const graded = await Promise.all(Array.from({ length: runs }, async () => grade(c, await system.answer(c.input))));
    results.push({ id: c.id, mustPass: c.mustPass, pass: graded.every(g => g.pass), why: graded.flatMap(g => g.why) });
  }
  const score = results.filter(r => r.pass).length / results.length;
  const blocked = score < 0.8 || results.some(r => r.mustPass && !r.pass);
  return { score, blocked, results };       // stored with prompt, model, index and code versions
}
gradercostgood forwatch out
code checkfree, instantcitations, numbers, format, forbidden contentmisses quality and tone
model rubrica model call per caserelevance, completeness, tonebiased toward long answers; grade the grader
human reviewexpensivecalibration, new failure typessample it; do not gate on it
the staff move
Make the eval the place where arguments end. When someone says the new prompt "feels better", the answer is to run it and read the newly failing list together. That turns taste into a reviewable diff.
EVALS IN DEPTH
turning "it seems better" into a number you can gate a release on
swipe the figure sideways, or tap expand for full screen
1/6
cases
Cases: start with 15 to 30 real questions covering the main intents, the known hard ones and the must-refuse ones. Write expected properties, not exact strings: "quotes the 50 naira fee", "cites fees.md", "refuses to promise a refund". M started with 16 cases; that is enough to begin.