Keeping a copy of an expensive result so the next request is fast and cheap. The hard part is knowing when the copy is out of date.
You have done this if
You cached embeddings for repeated queries, or cached model answers for identical prompts.
Say it in a review
We cache retrieval results with a short TTL and invalidate on re-index, because a stale answer looks exactly like a right one.
On the AI Application map Retrieval, LLM Gateway