New here? Start this series at Part 1: Logging the answer tells you almost nothing
- Part 1: Logging the answer tells you almost nothing
- Part 2: Your AI feature has unit economics whether you measured them or not
- Part 3: Real users ask questions your test set never imagined
- Part 4: You cannot roll back a prompt you never versioned
- Part 5: Nothing broke, and the system is still getting worse
The feature shipped in March and everyone was pleased. In June, finance asked why the model line had grown eleven times while revenue had not moved, and nobody in the room could say which product surface was responsible.
That conversation is avoidable, but only before launch.
Budget the workflow, not the call
Per-token pricing invites you to reason about one call. Users never experience one call. They experience a workflow that fans out into a retrieval query, three model calls, two tool round trips and a re-ask when validation fails.
So the unit is cost per successful task. Successful is load-bearing: a run that ends in a retry loop still burns tokens, and a metric that counts only completions will understate the real number. Same for latency. The p95 on your model call is not the number the user feels.
Context discipline is the biggest lever most teams have not pulled. Prompts accumulate. An example here, a policy paragraph there, full conversation history because summarising was harder. Nobody deletes anything, because deletion might hurt quality and quality is hard to measure. Six months later, half the spend is re-sending context that changes no answers.
Set the target per workflow
Not every path deserves the same budget. Decide the tier explicitly and write it down:
| Workflow class | Latency target | Cost ceiling | Default approach |
|---|---|---|---|
| Inline assist | under 300 ms | fractions of a cent | small model, hard cache, no tools |
| Interactive answer | 2 to 5 s, streamed | cents | mid model, retrieval, capped tool loop |
| Deep task | seconds to minutes | tens of cents | strong model, show progress, checkpoint |
| Batch or offline | hours | lowest per unit | batch API, off-peak, retry cheaply |
The numbers are yours to set. The point is that somebody chose them before shipping, so a regression becomes a visible breach rather than a drift nobody owns.
The levers, in the order I reach for them
Route first. Most traffic in most products is low risk and does not need your best model. Classify the request, send easy ones down the cheap path, reserve the expensive path for what earns it.
Cache second. Exact-match caching catches more than teams expect and costs almost nothing. Semantic caching helps too, with a real failure mode: a near-miss returning a confidently wrong neighbour. Give it its own eval.
Batch third, wherever latency allows. Then degrade deliberately. Under load, a shorter answer from a smaller model beats a queue timeout, and you should decide which surfaces may degrade before you are inside the incident. Capacity planning is the unglamorous part: provider rate limits, queue depth, behaviour at three times normal load. Test the queue before your users do.
Attribution, or that June meeting
Tag every call with feature, tenant, model and outcome. Report spend along those dimensions weekly, to the team that can actually change it. Cost regressions should reach the same people quality regressions do.
Unit economics you did not choose are still unit economics. You just meet them later, in a meeting you did not schedule.
More on these topics
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Explainer · · 1 min read
Why autonomy needs budgets
Why autonomous systems need limits on cost, time, actions, retries, and risk.
Explainer · · 2 min read
Why tokenization quietly affects cost, limits, and reliability
How a low-level text detail quietly becomes a product constraint for cost, latency, and reliability.
Discussion