Part 8 · 1 chapters · ~12 min
Evals in Depth
An eval as a test suite for non-deterministic behaviour: cases written as expected properties, code graders before model rubrics, a threshold with must-pass cases, version-to-version diffs, growing the set from production failures with a held-out slice, and running it in CI, nightly and in mock mode.
9
Cases, graders, gates and diffs
code
// an eval harness in miniature
type Case = { id: string; input: string; tags: string[]; mustPass?: boolean;
expect: { cites?: string[]; contains?: string[]; absent?: string[]; refuses?: boolean } };
async function grade(c: Case, out: Answer): Promise<{ pass: boolean; why: string[] }> {
const why: string[] = [];
for (const s of c.expect.cites ?? []) if (!out.citations.includes(s)) why.push(`missing citation ${s}`);
for (const s of c.expect.contains ?? []) if (!out.text.includes(s)) why.push(`missing "${s}"`);
for (const s of c.expect.absent ?? []) if (out.text.includes(s)) why.push(`forbidden "${s}"`);
if (c.expect.refuses !== undefined && out.refused !== c.expect.refuses) why.push('refusal mismatch');
return { pass: why.length === 0, why };
}
async function runEval(cases: Case[], system: System, runs = 3) {
const results = [];
for (const c of cases) {
const graded = await Promise.all(Array.from({ length: runs }, async () => grade(c, await system.answer(c.input))));
results.push({ id: c.id, mustPass: c.mustPass, pass: graded.every(g => g.pass), why: graded.flatMap(g => g.why) });
}
const score = results.filter(r => r.pass).length / results.length;
const blocked = score < 0.8 || results.some(r => r.mustPass && !r.pass);
return { score, blocked, results }; // stored with prompt, model, index and code versions
}| grader | cost | good for | watch out |
|---|---|---|---|
| code check | free, instant | citations, numbers, format, forbidden content | misses quality and tone |
| model rubric | a model call per case | relevance, completeness, tone | biased toward long answers; grade the grader |
| human review | expensive | calibration, new failure types | sample it; do not gate on it |
the staff move
Make the eval the place where arguments end. When someone says the new prompt "feels better", the answer is to run it and read the newly failing list together. That turns taste into a reviewable diff.
EVALS IN DEPTH
turning "it seems better" into a number you can gate a release on
swipe the figure sideways, or tap expand for full screen
1/6
cases
Cases: start with 15 to 30 real questions covering the main intents, the known hard ones and the must-refuse ones. Write expected properties, not exact strings: "quotes the 50 naira fee", "cites fees.md", "refuses to promise a refund". M started with 16 cases; that is enough to begin.