Part 7 · 1 chapters · ~12 min

Context Engineering and Retrieval

What goes into a model's window and why it decides the answer: chunking along structure with metadata, embeddings and the index, top-k similarity against MMR measured on real data, hybrid keyword and vector search with reranking, the context budget, and four retrieval failure modes with their fixes.

8

What goes into the window

Most bad answers from a grounded assistant are retrieval failures, not model failures. Debug the context before you touch the prompt: log what was retrieved for each bad answer and ask whether a careful human could have answered from it.

code
// chunk along structure, keep metadata, retrieve with filters, rerank
type Chunk = { id: string; text: string; source: string; section: string; product: string; updatedAt: string };

function chunkMarkdown(doc: { source: string; body: string; product: string; updatedAt: string }): Chunk[] {
  return doc.body.split(/\n(?=#{1,3} )/).map((sec, n) => {
    const heading = sec.match(/^#{1,3} (.*)/)?.[1] ?? '';
    return { id: `${doc.source}#${n}`, text: sec.trim(), source: doc.source, section: heading,
             product: doc.product, updatedAt: doc.updatedAt };
  });
}

async function retrieve(q: string, product: string, k = 5) {
  const qv = await embed(q);
  const dense = await index.query({ vector: qv, topK: 20, filter: { product } });     // meaning
  const sparse = await bm25.search(q, { topK: 20, filter: { product } });              // exact terms
  const merged = dedupe([...dense, ...sparse]);
  const scored = await rerank(q, merged);                                              // many in, few out
  return scored.slice(0, k);
}
symptomlikely causefirst fix
"I don't know" when the doc existschunk lost its heading; query uses different wordsheading in chunk text; hybrid search; query rewrite
quotes last year's feestale chunk ranks higherfilter or boost by updatedAt; delete superseded docs
right passage retrieved, wrong answertoo many passages; important one in the middlefewer, reranked passages; put the best first
confident answer, no citationprompt allows memory answersrequire citation ids; reject answers without them
your evidence
M is your example: phase 2 compared similarity retrieval with MMR on its own checks and chose similarity (95% against 75%). The story to tell is the method: you measured on your own documents instead of taking the popular default.
CONTEXT ENGINEERING AND RETRIEVAL
what goes into the window decides what comes out: chunking, embedding, retrieval, reranking and the budget
swipe the figure sideways, or tap expand for full screen
1/6
chunking
Chunking: split documents along their structure (headings, sections, list items), not every N characters, so each chunk is one idea with its heading attached. Add overlap only where sentences cross boundaries. Keep metadata (source, section, date, product) on every chunk: it powers filters and citations.