Tech #015•30 min read•19 August 2026 , Wednesday

MCP Was Built for Agents. Now It's Learning How to Scale.

The next challenge for Model Context Protocol isn't what agents can do — it's how MCP infrastructure survives when millions of requests start arriving.

Rajnish Kumar

Rajnish Kumar

Editor-in-Chief & Founder

MCP Was Built for Agents. Now It's Learning How to Scale. — Tech dispatch hero image

#01Why Does a Protocol for Talking to Tools Suddenly Care About Load Balancers?

On August 5, 2026, Google Cloud's engineering team published a detailed write-up titled "Scaling AI Agent Infrastructure with the MCP Stateless Updates." Nothing about that title sounds like it belongs in the same conversation as agents, tool calls, or language models — it sounds like a platform engineering post, the kind normally written about databases or message queues, not about the protocol that lets an AI assistant open a file, run a query, or call an API on your behalf. That mismatch is the actual story.

— the Model Context Protocol — was introduced by Anthropic in November 2024 as an open standard for one specific problem: giving an AI agent a common, predictable way to discover and call external tools, read resources, and use prompts, instead of every integration between every model and every tool being a bespoke, one-off wire format. Concretely, that means a client can call tools/list against any compliant server and get back a description of what it can do, then tools/call a specific one by name with structured arguments — the same three or four verbs, regardless of whether the server sits in front of a filesystem, a database, or a SaaS product's own API. It solved that problem well enough that within roughly eighteen months it went from a new specification to the default way a huge fraction of the agent ecosystem talks to the outside world — every major model provider ships an MCP client today, and directories of public MCP servers now list them in the thousands, covering everything from version control and cloud infrastructure to internal company tools nobody outside one engineering org will ever hear about. That's the part everyone already knows about MCP — what it lets an agent do.

What almost nobody was asking eighteen months ago is a much more boring question: what happens to that protocol once it stops running on one developer's laptop, talking to one tool server, for one conversation — and starts running inside Google Cloud, behind a real load balancer, serving millions of requests a day, across a fleet of servers that scale up and down constantly? A protocol built and battle-tested for the first shape of that problem doesn't automatically survive the second one; nothing about "works reliably for one developer at a desk" implies "works reliably for a platform vendor serving every customer's agents at once." The answer, it turns out, is that MCP inherited a very old, very well-understood distributed-systems problem, and spent its first two years not really having to notice.

That problem has a name, and it isn't specific to AI at all: statefulness. Any protocol that remembers something about a specific conversation, tied to a specific server process, runs into exactly the same set of failure modes the moment you try to run more than one copy of that server behind a load balancer — whether the protocol is MCP, a websocket chat server, or a 1990s telnet session. MCP didn't get a new problem in 2026. It got old enough, and popular enough, that the problem every stateful service eventually meets finally caught up with it.

That's also why this piece isn't going to read like the usual MCP news cycle — a new tool integration, a new client shipping support, a new capability an agent can now reach for. This is about the layer underneath all of that: the transport MCP actually runs on, and what had to change about it before a cloud platform could responsibly run it at the scale its own customers were already pushing it toward. The interesting question isn't "what can an agent do now that it couldn't before" — nothing here changes that at all. It's the one this piece keeps coming back to: why did a protocol for connecting agents to tools suddenly need to start caring, in detail, about round-robin routing, pod restarts, and where exactly a piece of state is allowed to live.

A protocol doesn't need to be about AI to inherit AI-scale traffic — it just needs to become popular enough that someone tries to run it the way real infrastructure actually runs.

#02The Old Mental Model: One Agent, One Session, One Server

To see why load balancers became a problem, it helps to be precise about how MCP originally worked — not the tool-calling part everyone pictures, but the plumbing underneath it. When an MCP client first connects to a server, the two sides perform an initialize/initialized handshake: the client announces its protocol version and capabilities, the server responds with its own, and — critically — the server may issue a session identifier back to the client in an Mcp-Session-Id HTTP header. From that point on, the client is required to include that same session ID on every subsequent request. The server, in turn, keeps whatever state belongs to that session — which tools have been discovered, what context has accumulated, what a long-running operation is currently doing — sitting in that one process's memory, keyed by that session ID.

This is a completely reasonable design if you only ever picture one agent talking to one server. An agent's relationship with a tool server is inherently conversational: it calls a tool, the response shapes the next call, maybe the server needs to ask a follow-up question mid-task. Holding that context in memory, tied to an open connection, is the simplest possible way to make that conversation feel continuous — no re-explaining who you are or what you were doing on every single message. It's the same intuition behind a database connection held open for a transaction, or a web framework's server-side session object: keep the state close to the code that's using it, and don't make the caller repeat itself.

It's also, worth saying plainly, the design almost anyone would reach for first. Handshake-once-and-remember is how a huge share of stateful network protocols have worked for decades — it's a genuinely good default when the thing on the other end of the connection is, in practice, always the same single process for the life of the interaction. MCP's original authors weren't solving a scaling problem when they specified sessions this way; they were solving a much narrower, much more immediate one — how does a client and a server agree on capabilities and keep a multi-step tool interaction coherent — and sessions are a perfectly sound answer to that narrower question. The scaling problem only exists once you ask a second question nobody had much reason to ask yet: what happens when "the server" is no longer a single, known process.

The mental model underneath that design is a straight, single-file line — one agent, going through one client, into exactly one server process, which is the sole owner of exactly one session:

Nothing about that diagram is wrong, exactly — it's just describing a single machine, or at most a fixed, known server a client always reconnects to. It has no opinion at all about what happens once "the server" stops being one process and becomes a fleet of interchangeable pods behind a , which is precisely the environment Google Cloud's post is describing.

The stateful model isn't a bug — it's a design built around exactly one server existing, quietly carried into a world where that stopped being true.

Agent
↓
MCP Client
↓
Stateful MCP Server (the session lives in *this* process's memory)
↓
Session

#03What Breaks When You Put That Behind a Real Load Balancer

Google's engineering team is specific about what actually goes wrong, and none of it is exotic — it's the standard list of problems every stateful service eventually runs into the moment it needs more than one instance. A standard round-robin load balancer's entire job is to spread incoming requests evenly across a pool of backend instances, and it does that job well precisely because it doesn't know or care what happened on the previous request. That's exactly the property a session-based protocol can't tolerate: if request one lands on pod A and gets assigned a session there, and request two — carrying that same Mcp-Session-Id — gets round-robined to pod B instead, pod B has never heard of that session. The result is a 400 Session Not Found error, not because anything crashed, but because the load balancer did precisely what it was designed to do.

None of this is a new problem the AI industry invented — it's the exact same trade-off web application servers were making twenty years ago, back when a PHP or Java session lived in one server's memory and load balancers needed "session affinity" cookies to keep working at all. The web mostly solved it by moving toward stateless request handling wherever possible, and pushing anything that genuinely needed to persist into an external store designed for exactly that job — a database row, a signed cookie, a cache. MCP arriving at the same conclusion isn't a coincidence or a sign the protocol got something wrong the first time; it's a sign the protocol is finally being run at the scale where that same, well-worn lesson becomes unavoidable again.

The conventional workaround for this, long before MCP existed, is sticky sessions — configuring the load balancer to route every request from a given client to the same backend instance for the life of that session, usually via a cookie or a consistent hash. It works, in the narrow sense that it stops the 400 errors. But it trades that fix for a worse, quieter problem: sticky routing actively fights the load balancer's actual purpose. A pod that happens to be "sticky-owned" by a disproportionate number of long-lived sessions can't shed that load onto idle neighbors, even while those neighbors sit underused. Autoscaling — spin up new pods under load, terminate them when load drops — becomes far less effective, because new pods can only pick up new sessions, never rebalance existing ones. You end up paying for elastic infrastructure while getting something closer to a fixed, unevenly loaded pool.

Then there's the failure mode that has nothing to do with load balancing at all: pod restarts. Rolling deployments, autoscale-down events, and ordinary crash recovery all do the same thing to a stateful server — the process ends, and every session it was holding in memory disappears with it, instantly and without warning. A client mid-conversation with that pod doesn't get a graceful handoff; it gets an error, because the state it depended on simply no longer exists anywhere. In a stateless HTTP service, restarting one instance behind a load balancer is invisible to callers by design — that's the entire point of running more than one instance. A stateful MCP server can't offer that guarantee at all.

This is where the problem stops being a rare edge case and becomes a routine operational fact. A platform team doesn't restart pods occasionally, as some unusual event to be worked around — they restart them constantly, as an intentional, healthy part of running a service: a rolling deploy every time code ships, an autoscale-down the instant traffic drops, a node getting cordoned for maintenance, a crashed process getting replaced automatically. Every one of those is meant to be background noise, invisible to whoever's actually calling the service. A stateful MCP server turns every single one of them into a visible, user-facing failure instead — which means the more diligently a platform team follows normal, healthy operational practice, the more often they break active agent conversations, an inversion of incentives no infrastructure should have to live with.

The other common answer — used by teams who can't accept sticky sessions' downsides but still need multiple instances — is to move session storage out of any single process and into a shared external store, typically , that every pod can read from and write to. This genuinely solves the "wrong pod" problem: any instance can look up any session, because the session no longer lives inside a specific process. But it solves it by adding a new network hop and a new piece of infrastructure to every single request — a round-trip to Redis to load the session before the actual work can even start, plus a cluster someone now has to provision, scale, secure, and keep available, since the whole system's correctness now depends on it too.

Put a number behind it and the scale of the problem gets easier to feel. Picture a platform running MCP servers for a product launch — traffic climbs from a quiet baseline to a sustained spike inside a few minutes, exactly the kind of load pattern autoscaling exists to absorb. New pods spin up on schedule, right on cue. But every one of them starts empty: no session has ever been assigned to it, because sessions were only ever handed out by the pods that existed before the spike. A round-robin balancer sends a share of traffic to those brand-new pods anyway, since that's its entire job — and a share of in-flight sessions immediately start failing with 400 Session Not Found, at the exact moment the system is under the most load and least able to absorb confused retries on top of it. The infrastructure did exactly what it was built to do. The protocol running on top of it just wasn't built to tolerate that.

That one sentence, from Google's own write-up, is really the entire chapter compressed into a single line — every failure mode above is a direct, mechanical consequence of it being true.

Every workaround for statelessness — sticky sessions, shared Redis stores — is really the same trade: pay in complexity, latency, or operational surface area for the privilege of pretending the server still remembers you.

“Standard round-robin load balancers do not know which container holds which in-memory session.”

#04The Stateless Core: What Actually Changed on July 28, 2026

The MCP specification revision published on July 28, 2026 takes a more direct approach than any of those workarounds: rather than finding a cleverer way to share session state across instances, it removes the session from the protocol's core entirely. The initialize/initialized handshake and the Mcp-Session-Id header — the two mechanisms that created a durable link between a client and one specific server process — are gone from the stateless core. There is no longer a setup step whose entire purpose was to be remembered later.

In their place, every request becomes self-describing. Protocol version, client identity, and capability information — everything that used to be exchanged once, during the handshake, and then implicitly relied on for the rest of the session — now travels inside a _meta field attached to every individual request. A server no longer needs to have met a client before to understand exactly what it's dealing with; each call carries its own complete introduction. Nothing about the interaction depends on which specific process handled the previous call, because nothing about the previous call is assumed to still be sitting in memory anywhere.

It also changes what debugging looks like when something does go wrong. A stateful bug report used to come with an implicit, awkward question attached — which specific pod was this session actually on, and is its in-memory state still inspectable, or did it already get recycled before anyone could look? Reproducing an issue often meant hoping traffic would get sticky-routed back to the same instance a second time. With every request self-describing and independent, that question stops being relevant at all: a request either succeeds or fails on its own terms, using only what it carried with it, and reproducing it is a matter of resending the same self-contained payload — to any instance, including one that's never seen this client before — rather than hunting down which specific process still remembers what happened.

That has a second, quieter consequence beyond fixing the routing problem: it changes what it costs to bring a new server instance online. Under the old model, a fresh pod was genuinely useless for any in-flight session the moment it started — it had no sessions, and no way to acquire one except by waiting for new clients to connect to it specifically. Under the stateless model, a fresh pod is fully capable of handling any request the instant its process finishes booting, including one that's a continuation of work a completely different, now-terminated pod started five minutes earlier. That's the property that makes Serverless deployment — scale-to-zero when idle, spin up on demand, tear back down after — a realistic option for MCP servers rather than a theoretical one; a platform like Cloud Run can start an instance from nothing and have it usefully absorb load immediately, because "immediately useful" no longer depends on that instance having built up any history first.

That single change is what makes the second diagram possible — not a smarter load balancer, not a faster shared cache, just the removal of the thing that made instances non-interchangeable in the first place:

Every request can now go to any healthy instance, because every request already contains everything that instance needs to handle it correctly. This is, on paper, an unremarkable diagram — it's how ordinary stateless HTTP services have been drawn for two decades. The interesting part isn't the diagram itself; it's that MCP needed a real specification change to be allowed to look like this at all.

Removing the handshake didn't make MCP simpler by accident — it made every server instance interchangeable on purpose.

┌→ MCP Server 1
Agent → LB ──┼→ MCP Server 2
└→ MCP Server 3

#05Routing Without Reading the Body

Statelessness solves who can handle a request. It doesn't, on its own, help a gateway or proxy sitting in front of that fleet make smart decisions about how to route or audit each one — and MCP's requests are JSON-RPC calls wrapped in an HTTP body, which means the only way to know what a request is actually doing used to be parsing that body first. A gateway that wanted to route tools/call requests differently from resources/list requests, rate-limit a specific tool, or log which tool was invoked for billing purposes had to unpack JSON on every single request just to make that decision — the deep packet inspection Google's post calls out directly as a cost the stateful model imposed on infrastructure.

It's not a coincidence that this specific piece of the update comes from a cloud platform vendor's own engineering team, rather than from a single-tenant tool author. A team running one MCP server for their own product can get away with an application-level router that already understands their own JSON shape intimately. A platform running MCP infrastructure on behalf of hundreds of different customers, each exposing entirely different tools, can't assume that intimacy — it needs to route, meter, and bill traffic it has never seen the specific shape of before, using the same generic mechanisms it already applies to every other kind of HTTP traffic crossing its edge. Headers are exactly the layer built for that: infrastructure that doesn't need to understand what a request means to still make a correct decision about where it should go.

Observability gets the same upgrade almost for free. A monitoring or logging pipeline built for generic HTTP traffic already knows how to key metrics off request headers cheaply — request-per-second by method, error rate by endpoint, the ordinary dashboards every service gets by default. Before this change, doing the equivalent for MCP traffic meant either building custom instrumentation that understood JSON-RPC specifically, or living without that visibility. With Mcp-Method and Mcp-Name sitting in plain headers, an operations team can get a genuinely useful first dashboard — call volume per tool, latency by method, error rate by specific tool name — out of infrastructure that has no idea what MCP is, using exactly the same approach it already uses for every other HTTP service in the fleet.

The July 28, 2026 specification addresses this with a small, deliberately boring change: three standardized HTTP headers that carry the information a router actually needs, without requiring it to touch the body at all. Mcp-Protocol-Version identifies which revision of the protocol the request is speaking. Mcp-Method carries the JSON-RPC method being invoked — tools/call, resources/list, and so on — directly as a header value. Mcp-Name goes one level more specific, naming the exact tool, prompt, or resource the request is targeting. A proxy, gateway, or load balancer can read all three the same way it already reads any other HTTP header, using infrastructure that has existed since long before MCP did.

Concretely, that's the difference between a router that has to parse this —

— and one that can make the identical routing decision just by reading Mcp-Method: tools/call and Mcp-Name: search_docs off the headers, before ever touching the body. It's a small mechanical change, but it's exactly the kind of change that lets MCP traffic be routed, rate-limited, and observed by completely generic HTTP infrastructure — the same gateways and proxies already routing everything else in a cloud environment, rather than something purpose-built to understand JSON-RPC.

If reading intent off a request without touching its body sounds familiar, it's the same instinct behind every layer of infrastructure that routes traffic before an application ever sees it — What Actually Happens When You Type a URL walks through exactly that chain, from DNS resolution to the load balancer, for an ordinary HTTP request; this chapter is the same idea applied one protocol later, to MCP's own request headers instead of a browser's.

Stateless routing needed one more piece: infrastructure that could tell requests apart without reading them — and that piece was always going to be headers, because that's what headers are for.

POST /mcp
Content-Type: application/json
{"jsonrpc":"2.0","id":7,"method":"tools/call","params":{"name":"search_docs","arguments":{"query":"stateless mcp"}}}

#06The Hard Part: How Do You Have a Conversation Without a Session?

Everything so far solves the easy half of the problem — a request that arrives, gets handled, and returns a response in one shot. But MCP was never only that. Sometimes a server genuinely needs to pause mid-task and ask the client something: which of three ambiguously-named files did you mean, confirm before this tool takes a destructive action, or here's a parameter I'm missing that only the user can supply. In the old, stateful world, this was almost trivial — the server just held the connection open and waited for the answer, because it owned that connection and nothing else needed to know the conversation was paused. Take the session away, and that trick stops working: there's no guarantee the answer, when it eventually arrives, will land on the same server instance that asked the question in the first place.

The July 28, 2026 specification's answer to this is Multi Round-Trip Requests, or MRTR — proposal SEP-2322, one of the Specification Enhancement Proposals that make up MCP's governance process (modeled directly on Python's PEPs: a numbered, publicly discussed document that a change has to pass through before it becomes part of the spec). It's the single cleverest idea in the whole update, because it doesn't try to preserve the old "hold the connection open" behavior at all. Instead of waiting, a server that needs input mid-call simply returns a normal, complete response: an InputRequiredResult, carrying a serialized requestState blob that fully describes exactly what the server needs to resume — which question was asked, what progress had already been made, everything necessary to pick the task back up from that exact point. The connection closes. Nothing is held open. Nothing is remembered anywhere.

The client goes and gets the actual answer — from the user, from the model, from wherever it needs to come from — and then reissues the call, echoing that same requestState value back as part of the new request. And because the state needed to resume the task is now traveling inside the request itself, rather than sitting in one particular process's memory, that follow-up request can land on any healthy instance behind the load balancer, not specifically the one that asked the original question. Instance 2 can finish a conversation instance 1 started, and neither side has to know or care that a handoff even happened.

Walk through a concrete case and the mechanism stops feeling abstract. Say an agent calls a delete_stale_branches tool against a Git hosting MCP server, and the server finds 400 branches matching the criteria — enough that it genuinely shouldn't proceed without a human confirming first. Instead of holding that request open while it waits for a "yes," the server responds immediately with an InputRequiredResult: a question ("confirm deletion of 400 branches?") plus a requestState blob encoding exactly where it got to — which branches it matched, which repository, which step of the operation it was mid-way through. That response travels all the way back up to the user, who confirms. The client then fires a new request, carrying the user's answer and that same requestState value. Nothing requires that new request to land back on the exact pod that made the original match — any instance can decode the requestState, see precisely what was being confirmed, and carry out the deletion, because everything it needs to act correctly was sitting inside the request the whole time. Concretely, the resumed call looks roughly like this:

No cookie, no session header, no dependency on which pod happens to answer — the entire "what were we in the middle of" question is answered by that one opaque field, which the server that issued it can decode and any server that understands the format can verify.

It also changes what "the user took too long to respond" actually means for the server. In the old model, a slow answer meant a connection sat open, consuming a slot on whichever process was holding it, for however long the user took — minutes, in the worst case, tying up a resource the whole time for no reason other than waiting. In the stateless model, there's nothing to tie up in the meantime at all: the server answered, in full, the moment it needed input, and simply isn't involved again until a new request shows up carrying the response. A requestState value can also be given its own expiry, so a confirmation nobody ever answers doesn't linger as valid forever — it just stops being a request an instance will honor once enough time has passed, the same ordinary kind of TTL already doing work elsewhere in this update.

MCP pairs this with a second, related mechanism for a slightly different problem: genuinely long-running work, the kind that takes ten seconds to a minute rather than needing a mid-call answer. For that case, a server returns a taskId immediately instead of blocking the connection at all, processes the work asynchronously, and lets the client check in later using standard tasks/get and tasks/update calls — polling instead of holding a connection open, for exactly the same reason MRTR avoids it: an open connection is a tether to one specific process, and stateless infrastructure can't offer any process that guarantee anymore. The distinction between the two mechanisms is really about why the connection would otherwise stay open — MRTR exists because the server is waiting on someone else to answer a question; the Tasks extension exists because the server itself just needs more time to finish.

The state didn't disappear when the session did — it just stopped living inside a process and started traveling inside the request that needs it.

POST /mcp
Mcp-Method: tools/call
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "delete_stale_branches",
"arguments": { "confirmed": true },
"requestState": "eyJvcCI6ImRlbGV0ZV9zdGFsZV9icmFuY2hlcyIsIm1hdGNoZWQiOjQwMCwic3RlcCI6ImNvbmZpcm0ifQ"
}
}

#07Caching Without Guesswork

One smaller piece of the same update is worth calling out on its own, because it follows the exact same philosophy in miniature: a ttlMs field, part of proposal SEP-2549, lets a server attach an explicit freshness window — in milliseconds — to the results it returns for things like resource or tool listings. A cacheScope value alongside it tells the client how broadly that freshness applies — whether this specific client can rely on the cached answer, or whether it's safe for any client talking to the server to reuse it.

The problem this replaces is a familiar one: without any explicit signal for "how long is this still valid," a client's only genuinely safe options are re-fetching constantly, just to check for changes, or keeping a long-lived SSE (Server-Sent Events) stream open purely to be notified the moment something does change — and a long-lived stream is, again, a connection tied to one specific server instance, the exact kind of dependency the rest of this update is working to remove. An explicit, numeric TTL turns "keep asking, or keep a connection open, just in case" into "trust this answer for the next 30,000 milliseconds and don't ask again until then" — a small change, but one more case of a stateless request/response exchange replacing something that used to require a standing connection.

It's a genuinely minor piece of the specification by line count, but it's worth noticing precisely because it's minor — it's the same underlying instinct as MRTR and the stateless core, just applied somewhere far less dramatic. Once a spec starts asking "does this really need a standing connection, or would an explicit value in the response do the same job," it tends to find more than one place to apply that question. Caching freshness was simply the smallest one.

Every long-lived connection this update removes has the same shape: something that only existed to answer a question a plain, boring timestamp could have answered instead.

#08What MCP Actually Needs to Remember

None of this means MCP servers can never have state — plenty of real tools genuinely need to remember something between one call and the next: a search that's paging through results, a multi-step workflow, a long file upload in progress. What changed is where that memory is allowed to live. It's no longer acceptable for it to live implicitly inside one server process's memory, tied to a session nobody but that process can see. It has to become something explicit and portable instead.

The pattern the specification points to is direct: if a server needs to carry state across calls, it mints an explicit handle from a tool call — a token, essentially — and hands it back to the caller. The model is then responsible for passing that same handle back as an argument the next time it needs to continue that piece of work. The state hasn't gone away; it's just no longer something the server silently remembers about a connection. It's a value the caller now explicitly holds and explicitly presents, the same way a database cursor, a pagination token, or a signed already work — state promoted from an implicit side effect of a connection into an ordinary piece of data any instance can read and act on.

A search tool paging through results makes the pattern concrete. The first call returns a page of matches plus an opaque cursor value; the model simply passes that cursor back as a normal argument on the next call, and the server — any instance of it — knows exactly where to resume:

Nothing about that call requires the server to remember the previous one issued that cursor. The cursor is the memory — encoded, portable, and exactly as meaningful to a server instance that has never seen this client before as it is to the one that generated it.

That reframing is really the entire chapter's answer to "does going stateless mean MCP forgets everything." It doesn't — it just refuses to let the protocol's own transport layer be the thing doing the remembering. Anything that genuinely needs to persist becomes the application's explicit responsibility, carried in plain data, instead of the protocol's implicit one, carried in a process's memory.

Handing state back to the client this way does come with an obligation a server didn't used to have to think about: that value is now sitting somewhere outside the server's own memory, which means a careless implementation could let a client tamper with it — inflate a page offset, forge a requestState it was never issued, replay one after it should have expired. The fix isn't complicated, but it is mandatory: treat every handle as untrusted input the moment it comes back, the same discipline already standard for a JWT or a signed cookie — sign or encrypt anything that carries information the client shouldn't be able to alter, and validate it fully on the way back in rather than trusting that a value merely being present means it's legitimate.

Stateless MCP didn't remove memory from the protocol — it evicted memory from the transport layer and handed the job to whoever actually needs it, as an explicit value instead of an implicit assumption.

{
"name": "search_repositories",
"arguments": {
"query": "stateless infrastructure",
"cursor": "eyJvZmZzZXQiOjIwLCJzb3J0IjoicmVsZXZhbmNlIn0"
}
}

#09Why Redis Session Storage Becomes Unnecessary — But Not Redis Itself

Put the last few chapters together and one specific piece of infrastructure quietly stops earning its keep: the shared Redis store teams stood up specifically to let every pod look up every session, the workaround described earlier for surviving a load balancer without sticky routing. That pattern worked by centralizing session state somewhere every instance could reach — but it was only ever necessary because session state existed somewhere that needed centralizing in the first place. Once the session itself is gone, rather than merely relocated, there's nothing left for that particular Redis layer to store, look up, or keep consistent. The dependency doesn't get optimized; it becomes structurally unnecessary, along with the network round-trip every request used to spend fetching from it.

It's worth being precise about what that claim does and doesn't cover, because it's easy to overstate. This is specifically about Redis-as-session-store for the MCP transport layer — the thing that existed purely to answer "which pod does this session belong to." A tool implementation behind an MCP server might have entirely legitimate reasons of its own to use Redis: caching an expensive downstream API response, rate-limiting a specific user, deduplicating work. None of that goes away, and none of it was ever what this update was targeting. The distinction is between Redis doing real application work, which is exactly what it's good at, and Redis being pressed into service purely to compensate for a protocol that couldn't yet run statelessly.

The operational cost of that session-store layer was never just the Redis cluster's own hosting bill, either — it was every other piece of complexity a shared external store drags along behind it. Someone has to size it for peak concurrent sessions, not average load, since a session-store cache miss doesn't degrade gracefully, it just breaks the request the same way a wrong-pod round-robin miss did. Someone has to think about what happens if that store itself becomes unavailable — a single shared dependency that every server instance now needs to reach on every request is exactly the kind of thing that turns a regional network blip into a total outage instead of a contained one. None of that is hypothetical; it's the ordinary cost of any shared, synchronous, on-the-critical-path cache, and it's a cost this update removes by making the dependency unnecessary rather than by making it faster or more reliable.

There's a genuinely close parallel in web authentication, and it's worth drawing out because it makes the pattern feel less MCP-specific. Server-side session stores were, for a long time, the default way a web application remembered that a given browser had already logged in — exactly the same "keep it in memory or in a shared store, keyed by an identifier the client presents" shape MCP just moved away from. -based authentication displaced a large share of that not by building a faster session store, but by making the session unnecessary in the first place: the token itself carries everything a server needs to verify who's asking, signed and self-contained, so any server can validate it without looking anything up anywhere. MCP's explicit handles and stateless _meta field are the same move, one layer further down the stack — state that used to require a lookup now arrives already proven.

Removing a dependency because it's no longer necessary is a very different achievement from removing it because it stopped working — MCP's session-store Redis falls into the first category, not the second.

#010Does Stateless MCP Really Solve the Whole Scalability Problem?

It would be a much weaker piece if it stopped there and called scalability solved, because statelessness answers exactly one layer of the problem — transport — and leaves several real ones standing. Making every server instance interchangeable does nothing to make an individual tool call fast. A tool that queries a slow downstream database, calls a rate-limited third-party API, or does genuinely heavy computation takes exactly as long after this update as it did before; stateless routing changes which instance can pick up that work, not how expensive the work itself is. Horizontal scalability buys you the ability to add more instances under load — it doesn't shrink the load a single request represents.

There's a second, more practical limit: the SDKs implementing all of this — TypeScript, Python, Go, and C# — are explicitly beta as of this specification. A huge number of MCP servers already deployed in production were built against the older, stateful assumptions, and moving them isn't automatic; it's a real migration, on the same order as any other breaking-ish protocol change, even with an explicit compatibility path for state that needs to survive as explicit handles. "The protocol now supports stateless deployment" and "the ecosystem has actually finished migrating to it" are two different claims, and only the first one is true today.

There's also a design cost this update doesn't erase, just relocates: someone still has to decide what belongs in an explicit handle and what doesn't, tool by tool, server by server. The stateful model made that decision for you, implicitly and often wrongly — everything just lived in the session, whether it needed durable, portable state or not. The stateless model forces every server author to actually think about it: is this piece of context safe and sensible to hand back to the client as an opaque token, or does it need to stay server-side for correctness or security reasons? That's a genuinely better question to be forced to answer, but it is a question — this update trades an easy default for a correct one, not for a free one.

There's a genuinely good answer to who this affects day to day, though, and it's a narrower group than the specification's own weight might suggest. Someone writing agent-side application code — deciding which tools to call, reasoning about a task, prompting a model — mostly shouldn't need to think about any of this at all; MCP client libraries absorb the handshake removal, the header changes, and the MRTR resume flow, the same way an application developer today doesn't hand-write TCP retransmission logic just because HTTP sits on top of it. This is squarely a server-author and platform-operator concern — the people actually running MCP servers behind real infrastructure, not the much larger population building on top of them.

It's also worth separating two kinds of "scale" that are easy to blur together. Horizontal scale — can I add more identical instances and handle more concurrent requests — is exactly what this update targets, and it targets it well. Vertical scale — how much work can a single request represent before it becomes a bottleneck regardless of how many instances you have — isn't touched by any of this at all. A tool call that fans out to twelve downstream services, or waits on a slow third-party API with no timeout, is precisely as expensive under the stateless model as it was under the stateful one. Nothing here makes a slow tool fast; it only makes sure a slow tool doesn't also require sticky routing to a specific, possibly overloaded pod.

Read that table as what it actually is: a real, well-designed answer to a real infrastructure problem, not a claim that every scaling question about AI agents has now been closed. The interesting scaling problems in this space were never really about whether a request could reach a server — they're about whether the work behind that request is fast, correct, and affordable once it gets there. This update clears the first obstacle out of the way cleanly. It was never going to be the last one.

Solving the transport layer's scalability problem is genuine progress, not the whole problem wearing a disguise — and it's worth being honest about which one just happened.

AspectStateful MCP (pre-July 28, 2026)Stateless MCP (July 28, 2026+)
Session trackingMcp-Session-Id, held in server memoryNone — every request is self-describing via _meta
Load balancingRequires sticky sessions / affinityPlain round-robin works
Pod restart / rolloutBreaks active sessionsInvisible to the client
Mid-call user inputConnection held openMRTR — InputRequiredResult + resumable requestState
Cross-instance stateShared Redis session storeExplicit handle, passed back as an argument
Deployment modelLong-lived processesWorks on serverless, scale-to-zero

#011What This Means If You're Building an MCP Server Today

None of the preceding chapters are purely academic — they translate fairly directly into decisions someone building or operating an MCP server actually has to make right now, while the ecosystem is still mid-migration. The first is the least exciting and the most important: don't reach for the stateless core just because it's new. If a server genuinely only ever runs as one process, for one team, with no load balancer in front of it, none of the failure modes in this piece apply to it yet, and adopting the new transport buys nothing but a beta SDK and an unfamiliar mental model. Statelessness is a solution to a specific, identifiable pain — multiple instances, a load balancer, a need for graceful restarts — not a default best practice to adopt on principle.

For a server that does need to scale horizontally, the real work isn't flipping a configuration flag — it's auditing what the server currently keeps in session memory and deciding, field by field, what actually needs to survive as an explicit handle versus what was only ever there because sessions made it easy to be lazy about. Not everything a stateful server accumulates during a session is meaningful state that needs to persist; some of it is genuinely disposable, recomputable on the next call for less cost than the complexity of carrying it forward as a token. Treating that audit as a design exercise, rather than a mechanical find-and-replace, is what separates a clean migration from one that just relocates the same bugs into a new shape.

It's also worth building the mid-call confirmation and long-running-task cases early rather than as an afterthought, even if a server doesn't need them on day one. MRTR and the Tasks extension are the two mechanisms most likely to be genuinely awkward to retrofit later, precisely because they change the shape of a server's response flow rather than just where state lives — a server built assuming every call either succeeds or fails in one round trip has more to restructure than one built with InputRequiredResult and taskId in mind from the start, even if those paths sit unused until they're needed.

And given the SDKs are beta, it's worth treating this as an active migration to track rather than a switch to flip once and forget — the same caution any team would apply to adopting a beta SDK for anything else load-bearing. Reading the specification directly, rather than relying on secondhand summaries (this piece included), matters more than usual while the surface area is still settling.

The teams that get the most out of this update won't be the ones who adopt statelessness fastest — they'll be the ones who correctly figure out which parts of their own state were ever worth keeping in the first place.

#012The Real Story Isn't the Agent Getting Smarter

Step back from the header names and the SEP numbers, and the more interesting shift is what this update reveals about where MCP actually is in its own lifecycle. Every headline about AI agents this year has been about capability — what a model can now reason through, what a tool can now let it do, how much autonomy it can be trusted with. This update isn't about any of that. It's about 400 Session Not Found, sticky routing, and whether a pod restart breaks a conversation — the exact same unglamorous distributed-systems concerns every piece of internet infrastructure eventually has to answer, whether it's serving AI agents or a checkout page.

That's not a smaller story than the capability headlines. It might be the more telling one. A protocol only has to solve horizontal scaling, graceful failover, and stateless routing once it's carrying enough real traffic that those problems stop being theoretical — and MCP crossing that threshold, quietly, in an engineering blog post rather than a product announcement, says more about how load-bearing it's already become than any new capability could.

Every piece of internet infrastructure that ends up mattering goes through some version of this same transition, usually without much fanfare. HTTP itself started as a way to fetch academic documents and had to be pulled, piece by piece, into something that could carry the entire commercial web. TCP/IP wasn't designed with cloud-scale load balancing in mind either; the infrastructure around it grew to make that possible once enough traffic demanded it. None of those transitions were about the protocols themselves becoming more interesting on paper — they were about the unglamorous work of making something that already worked also survive being depended on by everyone at once. MCP going stateless reads like the same story arriving for the AI agent ecosystem specifically, on a timeline compressed from decades into about two years.

MCP's interesting evolution isn't that agents are becoming smarter. It's that the infrastructure underneath them is being forced to become boring, scalable distributed software — round-robin load balancers, stateless request handling, graceful pod restarts, the same unglamorous discipline every serious piece of internet infrastructure eventually has to earn. That's not a downgrade in ambition. It's what it actually looks like when something stops being a demo and starts being infrastructure other people are allowed to depend on.

The question was never whether agents would get smarter — it was always going to be whether the plumbing underneath them could survive being taken seriously, and on July 28, 2026, MCP's answer became a lot more convincing.

Found this useful? Share it

Have a technical response or architectural perspective to share with the engineering desk?

Submit Engineering Feedback