MCP, the Model Context Protocol, is the standard that lets AI agents use real tools, so an agent can query a database, file a ticket or post to Slack. The first MCP server in a company takes an afternoon. Someone wires the GitHub server into an agent, the demo lands, and the pattern spreads, because the whole point of a standard protocol is that the second server is easier than the first. Six months later there are thirty. Official ones, community ones somebody found on a registry, a few built in-house for internal systems. Nobody decided to run thirty. Each one was individually reasonable.
If you run agents at work, this will feel familiar even if MCP itself is new to you. Secrets get copied into every agent because that was the fastest way to ship. The model reads pages of tool definitions before it does anything useful. And when someone finally asks who touched the customer database last month, the answer takes a week to assemble.
I have been building an MCP gateway, and nearly every feature in it exists because one of the problems below showed up first. None of them are protocol problems. MCP is doing exactly what it promised. The trouble lives in everything the protocol left to you.
The N×M problem moved instead of dying
MCP's pitch was arithmetic. Before, every agent needed custom glue for every tool, which meant N clients times M pieces of glue. After, everyone speaks one protocol and the work drops to N plus M. The arithmetic is real at the protocol layer, and it is why adoption was fast.
At the operational layer, the multiplication survives. Five agents connecting directly to thirty servers is a hundred and fifty connection configs and a hundred and fifty credential grants. When the database server rotates a key, someone edits five deployments. When security asks which agents can reach the CRM, the answer is a grep across repositories. The mesh MCP was supposed to kill is back, wearing a standard wire format.
A gateway restores the arithmetic. Agents talk only to the gateway, and the gateway talks to the servers. That does not mean one giant URL for everything. Each server keeps its own route on the gateway host, something like gateway.company.com/github and gateway.company.com/slack, so agents still address tools individually. What collapses to one is the control point. Config, credentials and policy live in one place instead of in every client that ever shipped.
Every connected server bills you in context before the first call
An MCP client loads the tool definitions of every connected server into the model's context. That is how the model learns the tools exist, and it is a cost that scales with connections rather than usage. You pay it whether or not a single tool gets called.
The numbers turn bad quickly. The GitHub server alone exposes roughly ninety tools. Add Slack, a database, a search server and a few internal ones, and the model reads tens of thousands of tokens of JSON schema before the user says a word. Anthropic's engineering blog worked through an agent whose tool definitions and intermediate results consumed 150,000 tokens on a task that, restructured, needed 2,000.
The quieter cost is accuracy. A model choosing between fifteen tools chooses well. A model choosing between three hundred, several of them named search or list_issues across different servers, chooses worse, and every misrouted call costs latency and money. The gateway is where you fix both. Give each agent a curated slice of the catalogue instead of the whole thing, rename collisions away, and serve tool discovery on demand rather than front-loading every schema.
Credentials multiply faster than servers
Each server authenticates its own way, a personal access token here, an OAuth app there, a service-account key for the internal ones. Engineers push back on this point, and fairly. Every server already has auth, and most keep their own logs too. But thirty working auth setups are the problem, because none of them talks to the others. Direct connections mean every client deployment holds every secret it might ever need, at the widest scope anyone asked for. The agent that only reads Jira tickets carries a token that can also close them, because provisioning a narrower one was extra work on a busy Tuesday.
That sprawl is what turns one leaked laptop into a bad month. It is also the part a gateway fixes most cleanly, because the pattern is fifteen years old. Clients authenticate once, to the gateway; the gateway holds the downstream credentials, injects them per request, scopes what each agent may reach, and rotation becomes one change instead of an archaeology project. API gateways have done this for HTTP services since before microservices had a name. Agents are a new kind of client, not a new problem.
Tool descriptions are text from strangers, fed straight to your model
Public registries list thousands of community MCP servers, and a tool description is not passive metadata. It goes into your model's context and shapes behaviour. Within months of the protocol shipping, researchers at Invariant Labs demonstrated tool poisoning, where instructions hidden inside a description quietly redirect the model or route sensitive data out through an innocent-looking parameter. The rug pull is the patient version of the same attack. The server behaves during review, gets approved, then changes its descriptions in a later update that nobody re-reads.
Direct connections give you nowhere to stand against any of this. A gateway is the natural chokepoint. It is the one place you can allowlist which servers are reachable at all, pin versions, hash tool descriptions and alert when a diff appears, and keep the full log of which agent called which tool with which arguments. You will want that log for the postmortem either way.
Compliance is where direct connections finally collapse
Sooner or later an auditor, a security review or a customer questionnaire asks the same three questions. Which agents can touch customer data. Who approved that access. What did they actually do with it last quarter.
With thirty direct connections, the honest answers are a grep, a shrug and a log-collection project. With a gateway, they are a policy file and a query, because every call already flows through one place that records the agent, the tool, the arguments and the decision. Access reviews become reading a config instead of interviewing five teams. Offboarding an agent means deleting one entry, not hunting its credentials across deployments. The same log that debugs a bad Tuesday doubles as audit evidence.
This is the argument that gets a gateway funded. Token charts persuade engineers; the compliance story persuades the people who sign off on infrastructure, because it converts an unbounded liability into a component with an owner.
What it costs, and when to skip one
The fair accounting comes first. A gateway is a new component someone has to run, and it adds a hop to every call. It also concentrates risk. When it goes down, every agent goes down together, so it needs the availability engineering you would give any shared dependency.
There is a real case for skipping it. If you are one team with a handful of first-party servers and agents on trusted infrastructure, a config file and some discipline are enough, and a gateway is ceremony. The forcing function is a boundary, the day a second team starts reusing your servers, a third-party server crosses into the trust perimeter, or a compliance question arrives with a deadline. Cross any of those lines and the gateway is cheaper than the mess it replaces.
You also do not have to buy one. A thin proxy that speaks MCP on both sides and does three things, allowlisting, credential injection and request logging, covers most of the risk and is a modest build. Start there, and add quotas, tool filtering and description pinning as the server count grows.
Building one end to end has left me with strong opinions about the parts that were harder than they looked, like how to structure the policy layer, what to log so compliance answers are a query rather than a project, and how to keep the gateway from becoming the bottleneck everyone fears. I am happy to share how to approach building one and what it buys an organisation on compliance. If your company is heading past its tenth MCP server, that is roughly the point where this stops being theoretical, so say so in the comments or reach out, and tell me where your sprawl hurts first.
More on these topics
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Checklist · · 18 checks
When to bring in a compliance review
The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.
Checklist · · 29 checks
Security review checklist for an AI feature
What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.
Discussion