Kit · Glossary
Architecture glossary Name the thing you already did. 67 architecture terms in 7 groups; each one says what it means, how to recognise it in work you have done, and one line to use in a review or an interview.
No term matches that. Try other words, or search the whole site .
The ability to keep serving, perhaps in a reduced form, when a dependency fails, and to recover on its own afterwards. It is a property of the whole system, built from timeouts, retries, fallbacks and isolation.
You have done this if
You added a fallback model when the primary provider timed out, or made the app answer from cache while search was down.
Say it in a review
We designed for the model provider being unavailable: requests fail over to a second model and the UI says when answers are degraded.
Continuing to work correctly despite a component failing, usually through redundancy. Stricter than resilience, which allows degraded service.
You have done this if
You ran two replicas behind a load balancer so a node could die without users noticing.
Say it in a review
Losing one instance is a non-event; we run at least two in separate zones.
Turning off the expensive or broken part and keeping the rest working, and telling the user what they are not getting.
You have done this if
When retrieval failed, the assistant said it could not reach the documents instead of answering from memory.
Say it in a review
If retrieval is down we answer with 'I can't reach the sources right now', never an ungrounded guess.
also fail fast
A switch that stops calling a failing dependency for a while after repeated errors, so you fail fast instead of piling up waiting requests, then lets a few test calls through to see if it recovered.
You have done this if
You stopped sending requests to a model endpoint after a burst of 5xx errors and routed to a fallback for a minute.
Say it in a review
The gateway trips a breaker per provider; after five failures in thirty seconds we route around it.
Separate pools of capacity (threads, connections, quotas) so one noisy workload cannot starve the others. Named after the compartments in a ship's hull.
You have done this if
You gave batch summarisation its own model quota so it could not eat the interactive chat's rate limit.
Say it in a review
Batch and interactive traffic have separate quotas, so a big backfill can't slow down users.
also retries, exponential backoff
Trying a failed call again after a growing, slightly randomised delay, so many clients do not retry in lockstep and flatten the service they are waiting for.
You have done this if
You wrapped model and tool calls in retries that wait 1, 2, 4 seconds plus a random offset.
Say it in a review
Retries use exponential backoff with jitter and a cap, and only on errors that are safe to retry.
also deduplication, exactly once, duplicate requests
An operation that has the same effect whether it runs once or several times. It is what makes retries safe for anything with side effects.
You have done this if
You passed a request ID with every 'create ticket' call so a retried agent step did not open two tickets.
Say it in a review
Every tool call with side effects carries an idempotency key, so a retry can't send the email twice.
An explicit limit on how long each step may take, set so the steps add up to less than what the user will wait for.
You have done this if
You gave retrieval 800 ms and the model 20 seconds so the whole request stayed under the 30-second gateway limit.
Say it in a review
Each hop has a timeout and they sum to less than the user-facing deadline.
Letting a slow consumer signal upstream to slow down, by queueing with limits, rejecting with 429s or pausing intake, instead of collapsing.
You have done this if
Your ingestion queue stopped accepting new documents when the embedding workers fell behind.
Say it in a review
Ingestion applies backpressure: past a queue depth we return 429 and the producer slows down.
How much breaks, or can be damaged, when one thing goes wrong. You reduce it by scoping permissions, isolating tenants and limiting what one agent run can touch.
You have done this if
You limited an agent to one project's repositories so a bad run could not touch the rest.
Say it in a review
A compromised agent token can only reach the one workspace it was issued for.
also DR, RPO, RTO, backup and restore
Getting the service back after losing a region or a datastore. RPO is how much data you can afford to lose, RTO is how long you can be down.
You have done this if
You restored the vector index from a nightly snapshot and re-ran ingestion for the gap.
Say it in a review
Our RPO for the index is 24 hours because we can rebuild from source; RTO is about two hours.
How well the system handles more load by adding resources. Horizontal scaling adds more instances; vertical scaling makes one instance bigger.
You have done this if
You moved the API to stateless containers behind a load balancer so you could add replicas during a launch.
Say it in a review
The web and API tiers scale horizontally; the model side scales through provider quotas, which is the real ceiling.
Scaling up and back down automatically with demand, so you pay for what you use.
You have done this if
Your container app scaled to zero overnight and back up on the first request.
Say it in a review
The frontend scales to zero off hours; we accept a cold start of a couple of seconds.
A service keeps no session data in its own memory between requests, so any instance can serve any request. State lives in a database or cache.
You have done this if
You moved conversation history out of the server process into Redis so you could run more than one instance.
Say it in a review
API instances are stateless; conversation state lives in the memory store, keyed by session.
Keeping a copy of an expensive result so the next request is fast and cheap. The hard part is knowing when the copy is out of date.
You have done this if
You cached embeddings for repeated queries, or cached model answers for identical prompts.
Say it in a review
We cache retrieval results with a short TTL and invalidate on re-index, because a stale answer looks exactly like a right one.
Splitting data across several stores or indexes by a key, such as tenant or region, so no single one holds everything.
You have done this if
You built one vector index per business unit instead of one global index.
Say it in a review
Indexes are partitioned by tenant, which also gives us a hard boundary for access control.
also quotas, 429
Capping how many requests a user, team or app can make in a period, to protect capacity and spend.
You have done this if
You set per-team token limits on the model gateway.
Say it in a review
Every team has a token budget per minute at the gateway, and we return 429 with a retry-after when they hit it.
Latency is how long one request takes; throughput is how many requests you complete per second. Improving one often costs the other, as batching does.
You have done this if
You streamed tokens to the browser so the first words appeared quickly even though the full answer took longer.
Say it in a review
We optimise time to first token for chat and throughput for batch jobs; they get different settings.
One request triggering many downstream calls, such as several searches or several agents in parallel. It multiplies load and failure chances.
You have done this if
Your query planner split one question into four searches and merged the results.
Say it in a review
A single question can fan out to four retrievals, so we cap parallelism and set a shared deadline.
Estimating the load you will get and making sure quotas, instances and budgets are in place before it arrives.
You have done this if
You requested a higher provisioned throughput from the model provider ahead of a company-wide rollout.
Say it in a review
We sized for peak tokens per minute at launch and booked the provider quota two weeks ahead.
After a change, different parts of the system may disagree for a while but converge if no new changes arrive. The opposite is strong consistency, where every read sees the latest write.
You have done this if
A policy was updated in SharePoint and the assistant kept quoting the old version until the next index run.
Say it in a review
The index is eventually consistent with the source; freshness is about fifteen minutes and we show the document date.
When the network splits, a distributed store has to choose between staying consistent and staying available. Useful mainly as a reminder that you are making that choice.
You have done this if
You chose to keep serving slightly stale results during an outage instead of refusing all requests.
Say it in a review
During a partition we prefer availability for search and consistency for permissions.
Writing the intended action to your own database in the same transaction as the state change, then sending it from there, so you never record an action you did not send, or send one you did not record.
You have done this if
An approval and the email it triggers are saved together, and a worker sends emails from that table.
Say it in a review
Agent actions go through an outbox, so the approval and the side effect can't drift apart.
A long business process split into steps across services, where each step has a compensating action to undo it if a later step fails, instead of one big transaction.
You have done this if
If the booking failed after the payment, your workflow refunded the payment automatically.
Say it in a review
The multi-step agent task is a saga: every step that changes something has an undo.
Command Query Responsibility Segregation. Separate models for writing data and for reading it, often with a read store shaped for the queries.
You have done this if
You wrote documents to a source system and served search from a separate index built for reading.
Say it in a review
Writes go to the system of record; reads come from an index built for retrieval, which is CQRS in practice.
Streaming inserts, updates and deletes from a database as events, so downstream copies stay in step without full reloads.
You have done this if
You used Delta change feed so only changed documents were re-embedded.
Say it in a review
Ingestion is incremental from change data capture, so a re-embed touches only what changed.
Changing the shape of data or messages over time without breaking the readers that still expect the old shape.
You have done this if
You added a field to the tool response and kept old agents working by making it optional.
Say it in a review
Tool contracts evolve additively; removing a field needs a version bump.
Knowing where a piece of data came from and every step that transformed it on the way.
You have done this if
Each answer could be traced back to the chunk, the document and the ingestion run that produced it.
Say it in a review
Every citation resolves to a document version and the pipeline run that indexed it.
Keeping data, and the processing of it, inside a specific country or region, usually because of law or contract.
You have done this if
You pinned the model deployment and the index to an EU region for European users.
Say it in a review
Prompts, embeddings and logs stay in the EU region; we verified the provider doesn't route outside it.
also scoped permissions
Giving each user, service and agent only the permissions it needs for its job, and nothing more.
You have done this if
You gave the agent a read-only token for the CRM instead of the admin key the prototype used.
Say it in a review
Each tool gets a scoped credential; the agent can read tickets but only a human can close them.
No request is trusted because of where it comes from, such as the internal network. Every call is authenticated and authorised on its own.
You have done this if
Your internal tools still checked the user's token even though they were only reachable from inside the VNet.
Say it in a review
Being on the private network doesn't grant anything; every service validates the caller's token.
Several independent layers of protection, so one failing does not expose everything.
You have done this if
You had a WAF at the edge, auth at the gateway, permission filters in retrieval and output checks before the answer.
Say it in a review
Prompt injection is handled in layers: input isolation, scoped tools, approval for actions and output checks.
also OBO, token exchange, delegated access
A service swaps the user's token for a new one to call the next service as that user, so permissions and audit follow the real person instead of a shared service account.
You have done this if
Your agent called Graph with the signed-in user's delegated token, so it saw only that user's files.
Say it in a review
Tools are called on behalf of the user, so the agent can never see more than the person asking.
Keeping credentials in a vault rather than in code or config, injecting them at runtime, and replacing them on a schedule or after exposure.
You have done this if
You moved API keys from environment files into Key Vault and set a 90-day rotation.
Say it in a review
No secret lives in the repo or the prompt; keys come from the vault at runtime and rotate quarterly.
No single person, or agent, can both request and approve a sensitive action.
You have done this if
The agent drafts the refund, a human approves it, and a different system executes it.
Say it in a review
The agent proposes, a person approves, and the system of record executes; no one step can do all three.
A tamper-resistant record of who did what, when, and with which data, kept long enough to answer an investigation.
You have done this if
You logged every tool call with the user, the arguments and the result to an append-only store.
Say it in a review
Every tool invocation is audited with the user identity, inputs and outcome, retained for a year.
Labelling data by sensitivity (public, internal, confidential, restricted) so controls can follow the label.
You have done this if
Documents marked confidential were excluded from the general assistant's index.
Say it in a review
Retrieval respects sensitivity labels; restricted content is never indexed for the broad assistant.
A structured look at what an attacker might want, how they could get it through your design, and what stops them.
You have done this if
Before launch you listed the ways a malicious document could steer the agent and added a control for each.
Say it in a review
We threat-modelled the retrieval path; the top risk was indirect prompt injection from shared documents.
Being able to answer new questions about what the system is doing from the outside, using traces, metrics and logs, without shipping new code.
You have done this if
You could open one trace and see the prompt, the retrieved chunks, the tool calls and the latency of each step.
Say it in a review
Every request is one trace across gateway, retrieval, model and tools, so we can explain any answer.
also SLI, SLA, service level
An SLI is what you measure (say, the share of answers under 5 seconds). An SLO is your target for it (99 percent). An SLA is the promise to a customer, with consequences if missed.
You have done this if
You set a target that 95 percent of chat responses start streaming within 2 seconds and alerted on it.
Say it in a review
Our SLO is 99 percent of answers grounded and under 8 seconds over 28 days.
The amount of failure the SLO allows. While budget remains you ship; when it is spent you slow down and fix reliability.
You have done this if
You paused a prompt change rollout because groundedness had already dipped below target that month.
Say it in a review
We'd burned most of the error budget, so the new prompt waited until the eval regressions were fixed.
Written steps for handling a known kind of incident, so whoever is on call can act without guessing.
You have done this if
You wrote down how to switch to the fallback model and how to roll back a prompt.
Say it in a review
Provider outage and bad prompt release both have runbooks with a rollback in under five minutes.
Sending a small share of traffic to a new version first and comparing it with the old one before rolling out further.
You have done this if
You routed 5 percent of users to the new prompt and watched the eval scores before switching everyone.
Say it in a review
Prompt and model changes go out as a canary at 5 percent with automatic rollback on eval regressions.
Running the new version next to the old one and switching traffic over at once, keeping the old one ready to switch back.
You have done this if
You built a new vector index alongside the live one and flipped an alias when it was ready.
Say it in a review
Re-indexing is blue-green: we build the new index beside the old and swap the alias.
A runtime switch that turns behaviour on or off without a deploy, for a group of users or everyone.
You have done this if
You let one team try agent actions while everyone else kept read-only answers.
Say it in a review
Write actions are behind a flag per tenant, so we can turn them off in seconds.
Quality changing over time without any code change, because the data, the users' questions or the model behind an API changed.
You have done this if
Answer quality slipped after the provider updated the model version behind the same name.
Say it in a review
We pin model versions and track eval scores weekly to catch drift in the data or the model.
Manual, repetitive operational work that grows with the system and could be automated.
You have done this if
Someone re-ran the ingestion job by hand every Monday until you scheduled it.
Say it in a review
Re-indexing was toil; now it's triggered by change events and nobody touches it.
Components depend on each other through small, stable interfaces, so one can change or fail without dragging the others along.
You have done this if
You put the model behind your own gateway interface so swapping providers didn't touch application code.
Say it in a review
The app talks to our gateway, not to a vendor SDK, so changing provider is a config change.
The agreed shape and behaviour of an interface (inputs, outputs, errors, limits) that both sides build against.
You have done this if
You wrote a JSON schema for each tool and tested agents against it.
Say it in a review
Tools have typed contracts; the agent only sees fields the schema defines.
New versions keep working for clients built against old versions.
You have done this if
You added a new optional parameter to an MCP tool without breaking the agents already calling it.
Say it in a review
Tool changes are backward compatible by default; breaking changes ship as a new tool name.
On the AI Application map Tools
One front door for many APIs that handles authentication, rate limits, routing and logging in one place.
You have done this if
All agent traffic to internal systems went through API Management instead of direct calls.
Say it in a review
Internal APIs are reached through the gateway, which enforces auth, quotas and audit for every agent.
also pub sub, events
Components react to events (something happened) instead of calling each other directly, which decouples timing and lets new consumers join without changing the producer.
You have done this if
A 'document updated' event triggered re-indexing instead of a nightly batch.
Say it in a review
Ingestion listens for change events from the source systems, so freshness doesn't depend on a schedule.
Where messages go after they fail processing too many times, so they can be inspected and replayed instead of blocking the queue or vanishing.
You have done this if
Documents that failed to parse landed in a separate queue you checked each morning.
Say it in a review
Unparseable documents go to a dead-letter queue with the error, and we replay them after a fix.
A translation layer between your design and a legacy or external system, so its odd data model does not leak into yours.
You have done this if
You wrapped the old ERP in a small service that returned clean, typed objects to the agent.
Say it in a review
The agent never sees raw ERP responses; an adapter maps them to our own types.
An open protocol for exposing tools, resources and prompts to AI clients in a standard way, so one integration serves many agents.
You have done this if
You wrapped Jira and your internal search as MCP servers instead of writing custom glue per agent.
Say it in a review
Integrations are MCP servers behind a gateway, so each one is built once and governed in one place.
AI systems The newer vocabulary around models, retrieval and agents, mapped to the plain engineering underneath.
Making the model answer from supplied evidence (retrieved documents, tool results) rather than from what it absorbed in training, and showing that evidence.
You have done this if
Every answer cited the passages it used, and the app refused when no passage supported it.
Say it in a review
Answers are grounded: no supporting passage, no answer.
also RAG, retrieval
Fetching relevant information at question time and passing it to the model with the question.
You have done this if
You indexed policy documents and passed the top matches to the model for each question.
Say it in a review
It's RAG over the policy corpus, with hybrid retrieval and citations.
Checks around the model that block, change or route requests and responses, such as topic limits, PII filters or approval for risky actions.
You have done this if
You blocked answers that contained account numbers and sent refund actions to a human.
Say it in a review
Guardrails sit at the gateway for content and in the tool layer for actions.
Instructions hidden in content the model reads (a web page, an email, a document) that try to make it do something the user did not ask for.
You have done this if
You stopped treating retrieved text as instructions after a document told the assistant to ignore its rules.
Say it in a review
Retrieved content is data, never instructions, and tools with side effects need approval.
A repeatable way to score the system on a fixed set of real questions with known good answers, run on every change.
You have done this if
You kept 200 real questions with expected sources and ran them in CI before merging prompt changes.
Say it in a review
Every prompt, model or index change runs against the golden set and can't merge on a regression.
Using a model to grade outputs against a rubric, which scales evaluation but needs checking against human grades.
You have done this if
A second model scored answers for faithfulness and you spot-checked a sample each week.
Say it in a review
We use a judge model with a written rubric, calibrated against human labels every month.
The code around a model that turns it into an agent: the loop, tool execution, memory, limits, stop conditions and logging. Most of an agent's reliability lives here, not in the model.
You have done this if
You wrote the loop that calls the model, runs the tool it picks, feeds back the result and stops after ten steps.
Say it in a review
The model is the easy part; the harness owns budgets, tool permissions, retries and when to stop.
One internal endpoint in front of all model providers that handles routing, keys, quotas, logging, caching and fallbacks.
You have done this if
Teams called your gateway instead of OpenAI or Azure directly, and you could see spend per team.
Say it in a review
All model traffic goes through the gateway, which gives us one place for keys, cost and fallbacks.
also HITL, approval
A person reviews or approves at chosen points, usually before irreversible actions or low-confidence answers.
You have done this if
Drafted customer emails waited in a queue for an agent to approve before sending.
Say it in a review
The agent drafts; anything that leaves the company or moves money waits for a human.
Treating the model's input limit as a budget to spend on the most useful content, because more context can lower quality and always raises cost.
You have done this if
You trimmed retrieved chunks to the best five and summarised long conversation history.
Say it in a review
We budget the context: instructions, the top five passages and a rolling summary, not everything we have.
Having the model return data in a fixed schema such as JSON so code can use it reliably.
You have done this if
You asked for a JSON object with fields for intent and entities and validated it before acting.
Say it in a review
Anything a program consumes comes back as schema-validated JSON, never parsed from prose.
How several agents or steps are coordinated: a supervisor that delegates, a fixed pipeline, or peers following a protocol.
You have done this if
One agent planned and handed sub-tasks to a researcher and a writer agent, then reviewed their output.
Say it in a review
It's a supervisor pattern with typed hand-offs; we chose it over peers because ownership stays clear.