By the third AI feature, the pattern is visible. Each one has its own prompt file, its own redaction step, its own trace format, its own idea of how much conversation history to keep. Changing one PII rule takes three pull requests, and one team ships late because nobody told them.
Nothing here is broken. Four platform concerns simply got copied into product code before anyone gave them a home.
Two planes, two clock speeds
The data plane is the request path: retrieve, assemble context, call the model, run tools, return a response. It runs per request, it is latency-critical, and it should be as boring as possible.
The control plane decides how the data plane behaves. Which prompt version is live, which model tier this tenant gets, what policy applies, what gets sampled into the eval set, what share of traffic sees the new configuration. It changes between requests and holds still during them.
Once you draw that line, the layers sort themselves out:
| Layer | Control plane owns | Data plane does |
|---|---|---|
| Policy and permissions | rule definitions, per-tenant configuration | checks and redacts on the call |
| Memory and context | retention, scope, what may persist | assembles the context window |
| Evaluation | golden sets, thresholds, sampling rules | emits a trace in the shared schema |
| Rollout | prompt and model versions, traffic splits | reads the resolved configuration |
The right-hand column is deliberately thin. Data plane code should read decisions that were made somewhere else.
The test is whether behaviour can change without a deploy
It is uncomfortably easy to check.
Tighten a redaction rule for one tenant. Roll a prompt back to yesterday's version. Move ten per cent of traffic to a cheaper model for one workflow. If any of those needs a code change, a release train and a regression check across three services, the control plane exists only as a set of habits.
The observability half has the same test. Ask what percentage of AI requests across your whole product emit traces you can compare. If each feature invented its own span names, you can debug one feature at a time but you cannot answer "did quality drop this week", because there is no "this week" that spans features.
Evaluation belongs here too, instead of being bolted on at the end. If traces flow into a shared store with a shared schema, building a golden set is filtering. If they do not, it is a data engineering project every single time.
Closing thought
Nobody sets out to build a control plane. It accumulates as duplicated policy code, five trace formats and a prompt change that needs a release. Naming the layer early is much cheaper than extracting it from four features later.
You have a control plane either way. Is yours a system or a habit?
More on these topics
Deep dive · · 6 min read
Every question your AI readiness review asks was answered months ago
The last six days of a sixty-day series, and the pattern is that operability gets bought early or it does not get bought at all.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Architecture pattern · · 1 min read
Designing a Multi-Agent System for Real-World Use
How to design and run a multi-agent system in production, using the supervisor pattern, with its trade-offs and where human approval belongs.
Discussion