Part 7 · 1 chapters · ~12 min
Context Engineering and Retrieval
What goes into a model's window and why it decides the answer: chunking along structure with metadata, embeddings and the index, top-k similarity against MMR measured on real data, hybrid keyword and vector search with reranking, the context budget, and four retrieval failure modes with their fixes.
8
What goes into the window
Most bad answers from a grounded assistant are retrieval failures, not model failures. Debug the context before you touch the prompt: log what was retrieved for each bad answer and ask whether a careful human could have answered from it.
code
// chunk along structure, keep metadata, retrieve with filters, rerank
type Chunk = { id: string; text: string; source: string; section: string; product: string; updatedAt: string };
function chunkMarkdown(doc: { source: string; body: string; product: string; updatedAt: string }): Chunk[] {
return doc.body.split(/\n(?=#{1,3} )/).map((sec, n) => {
const heading = sec.match(/^#{1,3} (.*)/)?.[1] ?? '';
return { id: `${doc.source}#${n}`, text: sec.trim(), source: doc.source, section: heading,
product: doc.product, updatedAt: doc.updatedAt };
});
}
async function retrieve(q: string, product: string, k = 5) {
const qv = await embed(q);
const dense = await index.query({ vector: qv, topK: 20, filter: { product } }); // meaning
const sparse = await bm25.search(q, { topK: 20, filter: { product } }); // exact terms
const merged = dedupe([...dense, ...sparse]);
const scored = await rerank(q, merged); // many in, few out
return scored.slice(0, k);
}| symptom | likely cause | first fix |
|---|---|---|
| "I don't know" when the doc exists | chunk lost its heading; query uses different words | heading in chunk text; hybrid search; query rewrite |
| quotes last year's fee | stale chunk ranks higher | filter or boost by updatedAt; delete superseded docs |
| right passage retrieved, wrong answer | too many passages; important one in the middle | fewer, reranked passages; put the best first |
| confident answer, no citation | prompt allows memory answers | require citation ids; reject answers without them |
your evidence
M is your example: phase 2 compared similarity retrieval with MMR on its own checks and chose similarity (95% against 75%). The story to tell is the method: you measured on your own documents instead of taking the popular default.
CONTEXT ENGINEERING AND RETRIEVAL
what goes into the window decides what comes out: chunking, embedding, retrieval, reranking and the budget
swipe the figure sideways, or tap expand for full screen
1/6
chunking
Chunking: split documents along their structure (headings, sections, list items), not every N characters, so each chunk is one idea with its heading attached. Add overlap only where sentences cross boundaries. Keep metadata (source, section, date, product) on every chunk: it powers filters and citations.