An employee asks the assistant what the salary band is for a senior engineer. It answers, precisely, from a planning document in a folder they have never been able to open.

Nothing malfunctioned. Retrieval did exactly what it was built to do. Embeddings do not carry access control lists, and the index had no opinion about who was asking.

A source registry, written before the index exists

Most teams can list their sources. Very few can name an owner for each one.

Before anything gets embedded, every source needs a row: who owns it, what trust level it carries, how often it syncs, and whether it is authoritative or merely informative. A published policy page and a two-year-old draft in a wiki are not equal evidence, and the retriever will treat them as equal unless something tells it otherwise.

Trust level should be a field you retrieve on rather than a comment in a design doc. When two chunks contradict each other, the system needs a rule, and "whichever scored higher" is not a rule anyone would defend in a review.

Permissions belong inside the query

The UI hides what the user should not see. The retriever does not.

Permission filters have to run inside the retrieval call: tenant ID, group membership, document-level ACLs applied before scoring rather than after. Post-filtering a result set is the pattern that quietly leaks, because a chunk that was retrieved and then dropped still influenced ranking, and it still surfaces in any citation list somebody forgot to filter.

A useful test: can the user open every citation the answer gave them? If a link returns 403, retrieval already read something it should not have.

Freshness and deletion are correctness problems

Deletes are the failure that embarrasses teams in public. Someone removes a document from the source system, and it stays answerable for another six weeks because the sync job only handles creates and updates.

Three numbers deserve first-class metric status: index lag against source, deletion propagation time, and the count of indexed documents whose source no longer resolves. Lineage makes those numbers answerable. Every chunk should carry its source ID, version and ingestion timestamp, so "why did it say that" resolves in one query instead of an afternoon.

The system also needs permission to say no. When the only sources covering a question are stale beyond the freshness budget, or the user's permissions filter the candidate set down to nothing, refusing is the correct answer. Answering from whatever survived the filter is how you get a confident response built on the one document nobody thought to restrict.

Governance can look like the boring stage before the interesting one, but it is what makes retrieval trustworthy enough to ship.

What would your retriever do tomorrow if someone revoked access to a folder it has already indexed?

Checklist · · 12 checks

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.