The first time you debug a genuinely bad agent run, you go looking for the step where it went wrong. You read the trace from the top expecting an obvious blunder: a hallucinated fact, a malformed call, a moment where it loses the plot.
Often there is no such step. Every action is defensible on its own terms. It re-read the file because the previous result was ambiguous. It re-ran the query because the first came back empty. It re-planned because the plan had failed. Forty steps later it has spent real money, touched real systems, and achieved nothing anyone asked for.
That is the characteristic agent failure: locally sensible, globally absurd. It is also why the reflex that serves you well on single-call systems, tighten the prompt, does almost nothing here. There is no bad sentence to rewrite. The badness lives in the sequence.
Six days of a sixty-day series went into agentic foundations, and read together the through-line is sharper than any one of them stated: an agent is not a model that acts, it is a trajectory, and every control that reliably governs a trajectory is enforced from outside the model.
A goal is a specification, and it gets read literally
Handing an agent a goal feels like delegation. It is closer to writing a contract, executed by something with a compiler's literalism and no interest in what you meant.
Underspecify it and the agent does not quietly fill the gap with your intent. It fills it with whatever best satisfies the words you used. "Reduce open ticket count" is satisfied by closing tickets without resolving them. "Maximise response quality" is satisfied by burning thirty tool calls on a question that needed one lookup.
The review exercise worth stealing takes five minutes. Give the goal specification to a colleague and ask what a malicious but compliant reading would do. Not what it should do: what it could do while honouring every word. That reading is available to the agent too, and it is the one you meet once real inputs stop resembling your test cases.
The practical bar is that success must be machine-checkable. If the only way to tell is a human squinting at the output, you have not written a goal. You have delegated a feeling.
The agency is in the cycle, not in the model
Strip the mystique and an agent is a loop: observe, plan, act, check, decide, repeat. A single LLM call is prompt in, text out, finished. An agent looks at the world, acts, sees what changed, and revises.
That decomposition turns "the agent is behaving strangely" into four inspectable questions. Sound reasoning producing wrong actions usually means the observation step is feeding it stale state. A sensible plan with broken execution points at the tool layer. Failures nobody noticed mean the check station does not exist, which is startlingly common: plenty of production agents fire an action and assume it worked. The same mistake repeating means results are not feeding back into the plan, leaving a script with retries wearing an agent's clothes.
None of those four are fixed by a better prompt. They are fixed by changing what the loop does between steps.
Every loop also needs three doors: done, blocked, and out of budget. An agent that cannot recognise it is stuck will keep paying for the privilege, and because each step looks reasonable in isolation, nothing in the trace flags it. The exits have to be enforced from outside the loop, not left to the agent's judgement.
Planning is cheap, and wandering is billed per token
Put two traces side by side, a planning agent and a non-planning one on the same task, and the difference is immediate. The non-planner acts, backtracks, re-fetches data it already had, and hits the critical dependency at step nine, where unwinding costs nine steps of spend. The planner met that dependency at step zero, on paper, for one model call.
A plan earns its keep through three properties, none of them eloquence. Steps that can be marked done or not done without interpretation. Dependencies stated explicitly, which surfaces the blocker before you have paid for the branches behind it. And revisability, because a plan the agent cannot update is a script marching into a world that stopped matching its assumptions.
The second payoff rarely gets mentioned: the plan is documentation of intent. Audit a run three weeks later because the bill spiked, and the trace tells you what happened while the plan tells you what the agent thought it was doing. The gap between them is usually where the bug is.
The trigger worth encoding: plan when actions cost money or are hard to reverse. For a single-step read, planning is pure latency.
Tools are where the risk stops being reputational
A model without tools can be wrong. A model with tools can be wrong and send the email, update the record, issue the refund, or delete the rows. That is the line where AI risk becomes operational. Treat the tool list as what it is: attack surface.
Four rules hold up under real traffic. Strict typed schemas rather than free text, because a tool accepting an arbitrary query string has no specified behaviour. Narrow beats general, so five specific tools like get_order_status and issue_refund_under_50 instead of one run_query. Read-only by default. And every call logged with arguments, result and context, because when something goes wrong the tool log is the investigation.
For each tool, ask the blast-radius question: what is the worst outcome if this is called with the worst possible arguments, at the worst moment, by an agent that has been confidently misled? Assume it might be. If the answer is unacceptable there are three fixes. Narrow the scope, require human approval, or make it reversible. "The prompt tells it to be careful" is not one of them.
The pattern to watch is accumulation. Tools get added under deadline, each individually justified, never revisited, and six months in nobody can list what the agent can do. A recurring twenty-minute review is the cheapest control here.
Memory is four different problems sharing one word
"Just add memory to the agent" is like "just add furniture to the house." Which room, and for what?
Working memory holds current task state and lives for minutes. Episodic memory records past runs and lives for months. Semantic memory holds durable facts about the user and domain. Procedural memory encodes the way things get done around here. Four words, four separate data-lifecycle problems.
Teams answer "what should it remember" easily, and the answer is always "everything." The unscheduled questions decide whether memory is an asset or a liability. What expires each class, given episodic memory grows without limit unless something removes it. When a stored fact turns out to be wrong, does the correction win at retrieval time or sit alongside the original with equal weight? That is where a lot of "the agent keeps insisting" behaviour comes from. And who can read each class, because episodic memory is a privacy surface that arrived without a privacy review.
Procedural memory is the most treacherous. When the underlying process changes, the agent keeps executing the old one competently. The failure looks exactly like correct behaviour, because it was correct, last quarter.
Reflection that changes nothing is a diary
There is a clean test for whether reflection is real. Run the task, let it reflect, run it again, and diff the behaviour. If run two is identical to run one, you paid for a paragraph.
Models are extremely good at sentences that sound like learning. "I should have been more thorough." "I should have verified the data before proceeding." These read as insight and change nothing, because they connect to nothing the next run will encounter. Real reflection has three stages and most implementations build only the first: name the step that failed, convert it into a concrete change such as an added precondition, and persist it where the next run will hit it.
Measure it as a system property rather than reading outputs and nodding. Track error rate on repeated task types over time. If that line is flat, the debrief is theatre you pay for on every run. Worth knowing, because the instinct when an agent underperforms is to add more reflection, which makes a broken loop more expensive, not more effective.
One question, asked six ways
Read the six together and they answer one question from six angles. Not how do I make the agent smarter. How does this thing stop, and what proves afterwards that it should have.
Stop conditions on the goal. Exit doors on the loop. A plan that surfaces the blocker before the spending starts. Blast radius caps on every tool. Expiry and correction paths on memory. A reflection budget with a measured payoff. Six topics, six versions of the same control.
That is the real difference between a demo agent and a production one, and it is not capability. The first run of an underspecified agent usually looks better than it has any right to, which is exactly what delays the discovery.
Autonomy is not the feature you are building. Bounded autonomy is, and the boundaries are the engineering.
Where each of these goes deeper
Each idea above is a standalone post on my blog, with the specifics and failure modes:
- Why agents need outcomes, boundaries, and stop conditions
- Why agent behavior is a loop, not a single prompt
- Why planning reduces wasted AI actions
- Why tool access is where AI becomes operational risk
- Why agent memory is not one database
- Why reflection should change the next action
They form the Agentic AI Foundations chapter of a sixty-day series on building AI that survives real users. If you would rather take it from the top, Start Here lays out all sixty posts in ten chapters.
More on these topics
Comparison · · 1 min read
LangGraph vs the OpenAI Agents SDK
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Discussion