New here? Start this series at Part 1: Autonomy is a runtime decision, not a model decision
- Part 1: Autonomy is a runtime decision, not a model decision
- Part 2: Your tool registry is an access control list wearing a different name
- Part 3: Four kinds of agent state, and only one of them is memory
- Part 4: A retried agent job is a second chance to send the same email
- Part 5: The 3am agent run that nobody is watching
Somewhere in your codebase there is a dictionary mapping tool names to functions, and every entry in it is offered to the model on every request. That dictionary is your permission model. Nobody wrote it down as one, nobody reviews it as one, and it very likely contains something like issue_refund.
The registry got built as a catalogue: what exists, what it does, what arguments it takes. Useful. It stops being sufficient the moment a second user with different rights sends a request.
The three bindings every entry needs
Identity. Which credential does this tool execute under? A single broad service token makes every call succeed and makes every side effect anonymous. When finance asks who approved a refund, "the agent" is not an answer. Tools that act on a user's behalf should carry that user's identity down to the system being touched, brokered at call time rather than baked into the process environment.
Scope. What can that credential read or change, expressed narrowly enough to be interesting? "Database access" is not a scope. "Read orders belonging to the requesting user, within this workspace" is.
Approval level. Some tools are free. Some are staged and need a human confirmation. Some are unavailable in this environment regardless of who is asking. This belongs in the registry entry, next to the schema, rather than scattered across prompt text.
Schemas sit alongside those three, and they are where drift creeps in. The description shown to the model is a promise about behaviour. When the implementation gains a parameter or quietly widens what it deletes and the description does not change, the model is now planning against a system that no longer exists. Test the schema and the permission scope in the same test, because a tool with a correct schema and the wrong credential is still a security bug.
Audit is what makes the other three checkable
Filtering the catalogue per user is the visible half of the work. The invisible half is whether you can reconstruct, months later, that a specific tool ran with specific arguments under a specific identity on behalf of a specific person, and that the scope in force at the time actually permitted it.
Without that record, permission scoping is an intention. With it, scoping becomes a claim you can verify, which means you can safely widen it. Registries that log properly end up being the ones teams are willing to add powerful tools to.
Pick the most destructive tool in your catalogue: whose credential does it run under, and could you prove that to an auditor without opening the source?
More on these topics
Checklist · · 29 checks
Security review checklist for an AI feature
What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.
Explainer · · 2 min read
Your agent should not be a superuser
Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.
Deep dive · · 5 min read
One MCP server is an integration. Thirty is why you need a gateway
Notes from building an MCP gateway. The protocol standardised how agents call tools, then stayed silent about credentials, context budgets and who may call what. That silence gets expensive as the servers multiply.
Discussion