Part 6 · 1 chapters · ~12 min
AI Products
A support assistant built properly: grounding in retrieved, cited sources with honest refusals, a classify-retrieve-generate-check pipeline that routes money requests to humans, a cost firewall in code, deterministic stubs for mock mode, an eval set as a release gate, and the production feedback loop.
7
Grounding, cost firewalls, deterministic stubs and eval as a release gate
a model is a dependency that can be confidently wrong
- Grounding: answer only from retrieved sources, cite them, and refuse when nothing relevant is found.
- The pipeline: classify, retrieve, generate, check, then answer, refuse or route to a human.
- A cost firewall enforced in code.
- Deterministic stubs so everything around the model can be tested.
- Eval as a release gate.
- In production: every bad answer is added to the eval set.
code
// the model client behind an interface: real, stubbed, and budgeted
interface Model { complete(req: { system: string; prompt: string; maxTokens: number }): Promise<{ text: string; tokens: number }> }
class StubModel implements Model { // mock mode: deterministic
constructor(private recorded: Map<string, string>) {}
async complete(r) { return { text: this.recorded.get(hash(r.prompt)) ?? 'I don\'t know.', tokens: 0 }; }
}
class BudgetedModel implements Model { // the cost firewall
constructor(private inner: Model, private budget: Budget) {}
async complete(r) {
if (!this.budget.allow(r.maxTokens)) throw new BudgetExceeded(); // caller degrades to search results
const out = await this.inner.complete({ ...r, maxTokens: Math.min(r.maxTokens, 800) });
this.budget.spend(out.tokens);
return out;
}
}
// eval gate in CI: any change to prompts/, retrieval/ or docs/ runs it
// npm run eval -- --set evals/support-v14.jsonl --threshold 0.93 → exit 1 below threshold| risk | control |
|---|---|
| hallucinated fee or rule | retrieval-only answers, number-in-source check, eval cases for every fee |
| prompt injection ("ignore your instructions") | no tools that move money; output checks; untrusted text kept as data, never as instructions |
| cost blow-up | per-request, per-user and per-day budgets; kill switch |
| silent regression after a model upgrade | eval gate on any model, prompt or retrieval change |
| PII leakage | redaction before prompts; approved providers; logs with PII handling |
| the provider is down or slow | timeout, breaker, fallback to search results (Books part 9) |
the course, complete
AI-native engineering comes down to the old disciplines, applied harder: precise specifications, real verification, clear boundaries, explicit controls and measured outcomes. The agents are new. What makes them safe and useful is not.
AI PRODUCTS: GROUNDED, BOUNDED, TESTED
a support assistant that answers from the company's documents, with a cost firewall, a deterministic stub and an eval gate
swipe the figure sideways, or tap expand for full screen
1/6
grounding
Grounding with retrieval: the assistant does not answer from the model's memory. The question is embedded, the most relevant passages are retrieved from the company's approved documents (product terms, fees, procedures), and the model is told to answer only from them and cite which passage each claim came from. If retrieval finds nothing relevant, the right answer is "I don't know; here is how to reach support".