Part 3 · 1 chapters · ~8 min
Test Generation and PR Critique
What agents test well and what humans must add, tests that cannot fail and mutation testing as the honest check, generating tests from the spec rather than the code, agent PR critique with verified findings, and human review that teaches and feeds the standing context.
4
What agents test well and miss; review as teaching
many tests is not the same as good tests
- Agents do well at spec cases, tables, boundaries and error paths.
- Humans add the implicit domain rules and the cross-module and concurrency cases.
- Tests that cannot fail do harm. Mutation testing is the honest check.
- Generate tests from the spec, not from the code.
- Agent critique brings different blind spots. Verify each finding.
- Human review teaches, and each lesson becomes standing context.
code
// a test that cannot fail (mocks the unit under test) versus one that can
// ✗
vi.mock('./fees', () => ({ fee: vi.fn(() => 2250) }));
test('fee', () => { expect(fee(150000, 'NG')).toBe(2250); }); // tests the mock
// ✓ derived from the spec: 1.5% capped at ₦2,000, integer kobo, rounding half up
test.each([
[100_000_00, 'NG', 1_500_00], // ₦100,000 → ₦1,500
[200_000_00, 'NG', 2_000_00], // cap
[1_33, 'NG', 2], // rounding: 1.995 kobo → 2
])('CLR-F-01 fee(%i, %s) = %i', (amount, country, expected) => {
expect(fee(amount, country)).toBe(expected);
});
// then: npx stryker run → any surviving mutant in fees.ts is an untested ruleTEST GENERATION AND PR CRITIQUE
what agents test well, what they miss, and review that teaches instead of only gatekeeping
swipe the figure sideways, or tap expand for full screen
1/6
what agents test well
What agents test well: enumerating cases listed in a spec, table-driven tests over inputs, boundary values (0, -1, max, empty, unicode), error paths of an API, fixtures and factories, snapshot-free assertions on rendered output via Testing Library queries. Fast and thorough on what is written down.