Someone picks 512 tokens with 50 of overlap because a blog post said so, and it quietly becomes permanent infrastructure. Nobody revisits it, because no number ever says it is wrong.

Chunk size looks like a tuning parameter, but it is really a decision about what a complete answer looks like for the questions your users actually ask.

Chunk to the shape of the question

A fact lookup and a procedure have different natural units, and one strategy cannot serve both well.

Question typeChunk toWhy
Specific fact lookup200 to 400 tokens, tight boundariesPrecision matters, and extra text dilutes the match
Procedure or how-toWhole section, headings preservedStep 4 is useless without steps 1 to 3
Comparison or synthesisSection plus a parent summaryThe model needs whole units it can compare
Policy or contract clauseClause boundary, never mid-sentenceA half-quoted clause is a legal problem, not a retrieval one

Split on structure first: headings, list boundaries, table boundaries. Fixed-size splitting with overlap is only the fallback for unstructured prose. And carry the heading path into the chunk text itself, so a chunk reading "must be approved within 5 business days" also carries "Expenses > International travel > Approval" and stops being ambiguous.

Rerank before you pack

Retrieve wide, keep narrow.

Pull fifty candidates from hybrid search and score each one against the full query with a cross-encoder. It is slower per candidate and much more accurate than vector similarity, which is why it runs over fifty documents instead of fifty million. The effect is the right chunk moving from rank 14 to rank 2, and position in the context window changes what the model leans on.

Add a diversity rule at the same step. Five chunks from one document look like strong corroboration and are really one source repeated.

The context window is a budget, not a container

Bigger windows made this problem less visible without making it any less real.

Every extra chunk costs tokens, adds latency, and gives the model another place to find something plausible but irrelevant. Set the budget explicitly: a chunk count, a token ceiling, and a minimum relevance score below which a chunk is dropped even when room remains. Order the survivors so the strongest evidence sits where attention is best, and label each block with its source ID and section so citations can be generated rather than invented.

Weak evidence deserves a weak answer

The failure that erodes trust fastest is a confident answer assembled from chunks that all scored 0.31.

Pass the scores into the generation step along with the text, and give the system a refusal path. Below the threshold, the honest response states what was searched, says nothing conclusive was found, and offers the closest sources instead of synthesising a paragraph out of near misses.

What relevance score would your system have to see before it says it does not know?

Try it: the chunking lab shows how chunk size, overlap and top k decide whether retrieval hands the model the whole answer.

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.