Part 6 · 1 chapters · ~12 min

AI Products

A support assistant built properly: grounding in retrieved, cited sources with honest refusals, a classify-retrieve-generate-check pipeline that routes money requests to humans, a cost firewall in code, deterministic stubs for mock mode, an eval set as a release gate, and the production feedback loop.

7

Grounding, cost firewalls, deterministic stubs and eval as a release gate

a model is a dependency that can be confidently wrong
  1. Grounding: answer only from retrieved sources, cite them, and refuse when nothing relevant is found.
  2. The pipeline: classify, retrieve, generate, check, then answer, refuse or route to a human.
  3. A cost firewall enforced in code.
  4. Deterministic stubs so everything around the model can be tested.
  5. Eval as a release gate.
  6. In production: every bad answer is added to the eval set.
code
// the model client behind an interface: real, stubbed, and budgeted
interface Model { complete(req: { system: string; prompt: string; maxTokens: number }): Promise<{ text: string; tokens: number }> }

class StubModel implements Model {                         // mock mode: deterministic
  constructor(private recorded: Map<string, string>) {}
  async complete(r) { return { text: this.recorded.get(hash(r.prompt)) ?? 'I don\'t know.', tokens: 0 }; }
}

class BudgetedModel implements Model {                     // the cost firewall
  constructor(private inner: Model, private budget: Budget) {}
  async complete(r) {
    if (!this.budget.allow(r.maxTokens)) throw new BudgetExceeded();   // caller degrades to search results
    const out = await this.inner.complete({ ...r, maxTokens: Math.min(r.maxTokens, 800) });
    this.budget.spend(out.tokens);
    return out;
  }
}

// eval gate in CI: any change to prompts/, retrieval/ or docs/ runs it
// npm run eval -- --set evals/support-v14.jsonl --threshold 0.93   → exit 1 below threshold
riskcontrol
hallucinated fee or ruleretrieval-only answers, number-in-source check, eval cases for every fee
prompt injection ("ignore your instructions")no tools that move money; output checks; untrusted text kept as data, never as instructions
cost blow-upper-request, per-user and per-day budgets; kill switch
silent regression after a model upgradeeval gate on any model, prompt or retrieval change
PII leakageredaction before prompts; approved providers; logs with PII handling
the provider is down or slowtimeout, breaker, fallback to search results (Books part 9)
the course, complete
AI-native engineering comes down to the old disciplines, applied harder: precise specifications, real verification, clear boundaries, explicit controls and measured outcomes. The agents are new. What makes them safe and useful is not.
AI PRODUCTS: GROUNDED, BOUNDED, TESTED
a support assistant that answers from the company's documents, with a cost firewall, a deterministic stub and an eval gate
swipe the figure sideways, or tap expand for full screen
1/6
grounding
Grounding with retrieval: the assistant does not answer from the model's memory. The question is embedded, the most relevant passages are retrieved from the company's approved documents (product terms, fees, procedures), and the model is told to answer only from them and cite which passage each claim came from. If retrieval finds nothing relevant, the right answer is "I don't know; here is how to reach support".