Most AI architecture reviews ask one question: does it do the thing? Everyone nods, it ships, and the next four questions arrive as incidents.
What the diagram should show
Four boundaries, which are the four days behind this one.
A gateway every model call passes through, so routing, cost and audit have one home. A stated pattern for each feature (chatbot, workflow, agent or RAG), chosen deliberately. A control plane holding policy, prompt versions, memory rules and eval configuration, separate from the request path that consumes them. And explicit human and tenant boundaries, with approvals carrying evidence and events carrying identity.
If a reviewer can point at each of those on the page and name its owner, the review can move on. If any of them is implied, that is where the first incident will start.
Every dependency needs a written degraded mode
Ask what happens when the model provider is slow, the vector index is stale, a tool API is down, or a tenant sends ten times its usual volume. Answer with what the user sees, in words you would put in the interface.
Degraded is not the same as broken. A grounded answer with a "sources may be out of date" banner is a good degraded mode. A confident answer built on an index that stopped updating on Tuesday is a bad one, and the difference is entirely in whether someone designed it. Write the degraded mode next to each arrow on the diagram, and the gaps become embarrassingly obvious.
Rollout and rollback are design decisions
A prompt change is a production change. It deserves a canary, a metric, and a reversal that does not require a release train.
That means versioning prompts, models and retrieval configuration as artefacts you can pin. It means shadow traffic before a cutover so you compare on real inputs instead of a demo set. And it means a rollback that is one action, because the moment you need it you will be reading a dashboard at an awkward hour.
The readiness review
Five questions, and a name attached to each answer.
What is the cost and latency budget for this workflow, and what enforces it? What quality signal would tell us this regressed, and who watches it? Which dependency, failing, produces the worst user experience, and what is the degraded mode? How is a bad rollout reversed, and how fast? Who is paged, and what can support see without an engineer?
Capacity belongs here too. Rate limits are shared across your product, so one enthusiastic new feature can starve an old one. Budget them per workflow before adoption teaches you the hard way.
Closing thought
Feature completeness is the easiest thing to review and the least predictive of how a system behaves under real load. The rest of the list takes an hour, and it is the hour that decides whether launch week is exciting or quiet.
Quiet is the goal.
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Architecture pattern · · 1 min read
Designing a Multi-Agent System for Real-World Use
How to design and run a multi-agent system in production, using the supervisor pattern, with its trade-offs and where human approval belongs.
Discussion