The ability to keep serving, perhaps in a reduced form, when a dependency fails, and to recover on its own afterwards. It is a property of the whole system, built from timeouts, retries, fallbacks and isolation.
You have done this if
You added a fallback model when the primary provider timed out, or made the app answer from cache while search was down.
Say it in a review
We designed for the model provider being unavailable: requests fail over to a second model and the UI says when answers are degraded.
On the AI Application map LLM Gateway, Model
Read Why production agents need recovery design · Every model call should go through something you own