Tech #016•30 min read•24 August 2026 , Monday

Your AI Agent Doesn't Need More Tools. It Needs to Know Which Tool to Use.

As agents gain access to hundreds of tools across dozens of MCP servers and APIs, the bottleneck stops being what an agent can do and becomes whether it can find, choose, and safely use the right capability at the right moment.

Rajnish Kumar

Rajnish Kumar

Editor-in-Chief & Founder

Your AI Agent Doesn't Need More Tools. It Needs to Know Which Tool to Use. — Tech dispatch hero image

#01When "What Can It Do" Stops Being the Interesting Question

Picture a fairly ordinary support agent, the kind teams are already wiring up this year. A user types one sentence — "find my last order" — and somewhere behind that sentence sits everything the agent has been connected to: a customer database, an order-management API, a shipping tracker, a refunds tool, an internal ticketing system, a handful of internal servers nobody outside platform engineering has fully inventoried. In a demo, that connection list is five or six tools long, and the agent picks correctly almost by default. In production, at a company that has spent a year plugging every internal system into its agent stack, that list looks nothing like the demo anymore.

The question worth sitting with isn't "can the agent technically reach an order-lookup tool somewhere in that stack" — of course it can, that connection was wired up months ago. It's a narrower, more mechanical one: out of everything it's connected to, how does the agent know *which* tool is the order lookup, *that* it's allowed to call it for this user, and *how* to call it without also triggering three other tools that happen to share the word "order" in their name? That's not a hypothetical edge case. It's the default condition of any agent that's been in production long enough to accumulate real integrations, and it's a genuinely different problem from the one most teams spent 2024 and 2025 solving — which was almost entirely about whether an agent could use a tool at all, not about what happens once it can use thousands of them at once. Concretely, for a single request, the graph behind it looks a lot more like this:

Adding a thousandth tool to an agent used to be the hard part; knowing which one of the thousand to reach for is the part nobody built for.

User

"Find my last order"

AI Agent

connects to

Tools

500

MCP Servers

50

Internal APIs

100

Databases

20

#02The Way Most Agents Still Register Tools

The overwhelming majority of agents running in production today register tools the same way the very first tool-using agent ever did: a developer writes out, by hand, a fixed list of function definitions and hands the whole list to the model on every single call. It's the natural thing to build first, because it's simple, inspectable, and completely under the developer's control — nothing is discovered, nothing is guessed at, every capability the agent has was deliberately put there by someone who wrote the code.

At six tools, or even sixty, this costs almost nothing: the definitions are small, the model has no trouble telling getOrder apart from refundPayment, and a developer can hold the entire list in their head. Concretely, that pattern looks like this:

Hard-coding a tool list isn't wrong — it's just a decision that was only ever cheap at the scale most teams started at, not the scale they end up at.

typescript
const tools = [
searchCustomer,
createInvoice,
sendEmail,
getOrder,
refundPayment,
updateTicket,
];
const response = await model.call({
messages,
tools, // all six definitions, on every request
});

#03From Six Tools to Six Hundred

The pattern breaks down the same way almost every "just hard-code the small list" pattern in software eventually breaks down: not at sixty, but somewhere on the way to six hundred, when the list stops being something one team owns and starts being the union of every integration every team has ever shipped.

Anthropic's own engineering team has published exactly what that growth curve costs in practice, and the numbers are concrete rather than hypothetical: a fairly modest five-server setup — GitHub, Slack, Sentry, Grafana, and Splunk, 58 tools between them — consumes roughly 55,000 tokens of context before the conversation has produced a single word of actual work. Internally, Anthropic has seen real setups where tool definitions alone ran to 134,000 tokens — close to half of a full context window spent describing what the agent *could* do, before it's done anything at all. Public directories make the scale of the underlying problem even more concrete: Smithery, one of the larger public MCP server directories, lists more than 7,000 servers on its own — and that's before counting the internal APIs, databases, and homegrown tools no public directory will ever know about. Once an agent's real tool surface looks anything like that, "hand the whole list to the model every time" stops being a simplification and starts being the thing quietly eating the context budget the actual task needed. Plotted out, the curve that gets a team there looks unremarkable right up until it isn't:

What starts as a six-tool list a developer wrote by hand ends, a few integrations later, as a catalog nobody on the team could recite from memory — and the hand-written pattern never actually changed to notice the difference.

10 Tools

100 Tools

1,000 Tools

10,000 Tools

#04What Actually Breaks: Tokens First, Then Accuracy

The token cost above is the easy half of the problem to see, because it shows up directly on a bill and in a context-window meter. The harder half is quieter and more dangerous: past a certain tool count, the model doesn't just get slower and more expensive — it starts *choosing wrong*, even when the correct tool is sitting right there in the list it was given.

Research benchmarking tool-selection accuracy against catalog size has found the drop-off is neither small nor gradual, and the numbers are worth reading slowly: at 740 tools, several models tested were choosing correctly *less than a fifth of the time* — not because the right tool wasn't available, but because it was buried in a list the model could no longer reason over reliably. Researchers point to three overlapping mechanisms behind that collapse. The first is a version of the well-documented "lost in the middle" effect: signal from the user's actual request gets diluted by the sheer volume of unrelated tool descriptions sitting in the same context, the same way a genuinely relevant fact gets harder for a model to use once it's buried in the middle of a long document instead of near the top or bottom. The second is tool collision — once a catalog has real breadth, many entries end up doing genuinely similar things (get_order, fetch_order, lookup_order_status), and the semantic boundary between them blurs for a model reasoning over all of them at once, the same way a person skimming sixty similarly-named spreadsheet tabs starts picking the wrong one. The third is a documented positional bias: tools sitting in the middle of a long list are measurably less likely to be selected correctly than ones near the start or end, regardless of how well-suited they actually are to the task.

A common first instinct is to fight this by organizing the catalog better — grouping tools into namespaced categories, prefixing names by system (github_create_issue, slack_send_message), writing tighter, more distinct descriptions. That's worth doing regardless, and it genuinely helps at the margins, but it doesn't change the shape of the underlying curve, because the model still has to read every definition in every category to know which category even applies. Better organization makes a crowded room easier to navigate; it doesn't make the room smaller. The actual fix has to reduce how much the model looks at *before* choosing, not just how neatly that material is arranged — which is exactly the retrieval step covered next.

None of this means more tools are bad — an agent genuinely needs broad reach to be useful across a real business. It means the naive way of granting that reach, dumping every available definition into every request, actively works against the model's ability to use any of it well. The fix researchers keep converging on is a retrieval step ahead of selection: one benchmark found that adding retrieval-based tool selection — searching for likely-relevant tools before presenting a shortlist, rather than presenting everything — more than tripled accuracy on the same tasks, from 13.62% to 43.13%, while cutting prompt tokens by over half in the process. That's the same idea the rest of this piece keeps returning to in different forms: don't hand the model everything and hope; hand it a way to ask for what it needs. One 2026 benchmark tracking accuracy as the number of available tools grows found roughly this shape:

Past a few hundred tools, the failure mode isn't the model getting confused occasionally — it's the model getting the right answer less often than a coin flip, for reasons that have nothing to do with the model's underlying intelligence.

Tools availableApprox. tokens spent on definitionsTypical selection accuracy
~50 tools~8K tokens84–95%
~200 tools~32K tokens41–83%
~740 tools~120K tokens0–20%

#05Two Different Fixes From the People Who Hit This First

Anthropic, running its own agents against real customer tool catalogs, published two separate fixes for two separate slices of this problem in 2026 — and it's worth being precise that they solve different things, because teams often reach for one when they actually need the other.

The first is the Tool Search Tool, and it attacks the token-and-accuracy problem directly at the point where tools get loaded into context in the first place. Instead of handing over every tool definition on every request, a developer marks the bulk of their catalog with defer_loading: true. Claude then starts a conversation seeing only the search tool itself plus whichever small set of critical, frequently-used tools were left non-deferred — not the full 500-tool catalog, just a handle it can use to go looking. When the model actually needs a capability, it searches, gets back references to matching tools, and only *those* get expanded into full definitions in its context. The traditional approach in Anthropic's own testing consumed roughly 77,000 tokens before any real work began; with the Tool Search Tool, the same setup started at roughly 8,700 tokens — an 85% reduction, preserving about 95% of the context window for the actual task. The accuracy effect was just as significant as the token effect: on Anthropic's own tool-selection evaluations, Claude Opus 4 improved from 49% to 74%, and Opus 4.5 from 79.5% to 88.1%, simply by not being shown hundreds of irrelevant tools it then had to reason past.

The second fix, code execution with MCP, targets a different cost entirely: not the tool *definitions* sitting in context, but the *intermediate results* passing through it. A traditional tool-calling agent that reads a two-hour meeting transcript from one system and attaches it to a record in another has to route that transcript through the model's context twice — once coming out of the read, once going into the write — which can burn an extra 50,000 tokens on a single, fairly ordinary handoff between two tools. Anthropic's fix reframes MCP servers as code APIs an agent can call from within a sandboxed execution environment, rather than as a flat menu of functions the model invokes one at a time. Tool definitions themselves get the same on-demand treatment, organized as a browsable filesystem (./servers/google-drive/, ./servers/salesforce/) that the agent explores and reads from only when it needs a specific one, rather than a wall of definitions loaded up front. On the representative document-handoff scenario Anthropic published, that combination — filesystem-based discovery plus in-sandbox execution — took a task like "read this document from Google Drive and attach it to a Salesforce lead" from roughly 150,000 tokens down to about 2,000: a 98.7% reduction, on exactly the kind of multi-tool handoff a support agent performs constantly.

One fix stops the model from drowning in tool descriptions it doesn't need yet; the other stops it from drowning in the data those tools hand back — and a production agent at real scale usually needs both, not either.

typescript
// Executed in a sandbox, not passed through the model's context
const transcript = (await gdrive.getDocument({
documentId: "abc123",
})).content;
await salesforce.updateRecord({
objectType: "SalesMeeting",
recordId: "00Q5f000001abcXYZ",
data: { Notes: transcript },
});

#06Zooming Out: Discovery Was Never Really an MCP Problem

Everything so far has been about what happens *inside* a single agent's context window, once it's already connected to its tools. That's genuinely useful, and it's also only half the story — because before an agent can defer-load a tool, search for one, or execute one inside a sandbox, something upstream has to have already answered a more basic question: how did the agent know that tool existed at all, and that it lived at this particular MCP server, on this particular domain, in the first place?

For most of MCP's life, the honest answer has been "a person told it to." Someone read a README, copied a server URL into a config file, and restarted the agent. That works fine for the handful of servers any one developer personally knows about, and it's exactly the same shape of problem the tool-list-in-code pattern has at small scale — completely fine until the number of things being manually wired up stops being small. MCP itself already tried to solve a narrower version of this once: the official MCP Registry, launched in preview by Anthropic together with maintainers from PulseMCP, Block, and GitHub on September 8, 2025, gave the ecosystem a single, open, searchable catalog of publicly available MCP servers — a real improvement over READMEs and word of mouth, but scoped specifically to MCP servers, and specifically to servers, not to skills, not to agents another agent might want to delegate to, not to plain HTTP APIs described in OpenAPI. It solved discovery for one protocol. It didn't solve discovery for the much bigger, much messier space of "everything an agent might need to find," because MCP was never the only shape a callable resource comes in.

The MCP Registry proved discovery-as-infrastructure works for one protocol — which is exactly what made the gap for everything outside that protocol impossible to ignore.

#07This Isn't the Web's First Discovery Problem

Widen the lens past AI agents entirely and "how does a consumer find something a publisher put out there, without a person manually wiring the two together" turns out to be one of the oldest solved problems on the internet — just solved separately, more than once, for different kinds of resource. A browser doesn't ship with a hard-coded list of every website's IP address; it asks to resolve a name it's never seen before, at the moment it's needed, trusting the answer because the domain hierarchy itself is the trust anchor. A JavaScript project doesn't vendor in a copy of every package it might ever need; it asks a registry like npm for a name, gets back a specific version, and trusts the package because the registry — not the individual publisher — is the thing every consumer already trusts by default. Neither system required the person building the browser or the build tool to have personally heard of every domain or every package in advance. That's the exact property agent tooling has been missing until this year.

The parallel is close enough to be genuinely instructive about what ARD is and isn't trying to be. DNS and npm both settled on a layered answer: a decentralized publishing side, anyone can stand up a domain or push a package, paired with a small number of trusted resolution points a consumer actually queries. Neither system has one single global authority deciding what's good — DNS has root servers and delegated zones, npm has its own registry plus private mirrors enterprises run internally — and neither needed one, because trust in both systems is anchored to *identity* (this domain, this package name, this publisher) rather than to a central authority vouching for quality. ARD's catalogs-plus-registries split, and Marsman's explicit "many discovery services, not one global catalog" framing, is the same layered shape arriving for agent tooling: publish where you already have identity (your own domain), let registries built for different audiences do the aggregating, and keep quality and ranking as a separate, unsettled question — exactly as it still is for both DNS results and npm search results today.

None of this means ARD's problem is already solved just because DNS and npm exist — an agent catalog carries genuinely new requirements neither of those systems had to deal with, chiefly the cryptographic trust-manifest layer covered later in this piece, since a DNS record or a package name was never expected to prove what it's *capable of doing* on a caller's behalf, only where it lives. But it does mean the core architectural bet ARD is making — decentralized publishing, pluralistic resolution, identity-anchored trust rather than centralized curation — is a bet the web has already placed twice before, at enormous scale, and won.

ARD didn't have to invent a new shape for "let anyone publish, let anyone find it, trust the identity rather than a gatekeeper" — it borrowed one the web already proved works, at DNS and npm's kind of scale, decades before agents needed it.

SystemWhat gets publishedWhat resolves itTrust anchor
DNSA domain name → IP mappingRecursive resolvers, root/TLD serversDomain registration
npm / PyPIA package name → codeThe registry's own search/installRegistry account + package name
ARD`ai-catalog.json` → tools, agents, APIsRegistries indexing many catalogsDomain ownership

#08The Agentic Resource Discovery Specification

That gap is precisely what the Agentic Resource Discovery specification, announced June 17, 2026, is built to close — and it's worth being specific about the fact that it's a joint effort, not a single vendor's proposal. Google published the initial specification alongside Microsoft, with GitHub, Hugging Face, Cisco, Databricks, GoDaddy, NVIDIA, Salesforce, ServiceNow, and Snowflake all signed on as contributing partners — a wider coalition than MCP itself launched with, reflecting how much broader the problem is than "one company's tools." Srinivas Krishnan, the Google Cloud engineer who co-authored the spec, has been characteristically blunt about how hard the underlying problem actually is — worth keeping in mind through the mechanics below, and worth closing this chapter in his own words.

The specification itself is built from two primitives, deliberately kept small. A catalog is a machine-readable ai-catalog.json file an organization publishes at a well-known path on its own domain, describing the capabilities it's making available — MCP servers, A2A agents, OpenAPI-described tools, or even nested catalogs pointing further downstream. Because the catalog lives on a domain the organization actually owns, that ownership itself becomes the first layer of identity: an agent resolving mcp.example.com's catalog gets a cryptographically grounded reason to believe it's actually hearing from example.com, not from something impersonating it. A registry, in turn, is what crawls and indexes those catalogs across many domains, so that an agent can search by *task intent* — "what can handle a refund request" — instead of needing to already know which of fifty companies happens to expose the right capability. Put together, the specification describes a fairly linear four-step flow: an organization publishes its catalog; an agent discovers it, either through a registry search or by resolving a domain directly; the agent cryptographically verifies the catalog's trust metadata before going any further; and only then does it connect at runtime, using whatever native protocol that specific resource actually speaks — MCP, A2A, or a plain OpenAPI call. Visually, that fan-out looks like this:

It's important to be precise about what the specification deliberately doesn't do, because it would be easy to read it as one more competing protocol trying to replace MCP. Its own documentation is explicit on this point: ARD "is not MCP, A2A, Skills, AI Catalog, or an API runtime, and it does not replace them" — it operates entirely *before* invocation, as the layer that helps an agent locate the right resource, after which the actual call still happens over whichever native protocol that resource speaks. Microsoft's Jennifer Marsman made a related point about how the spec expects to actually be deployed in practice, one that's easy to miss if you picture a single master catalog somewhere: "The goal isn't a single global catalog of every resource. There will be many discovery services, each defined by what it indexes, whom it serves, and how it ranks." The specification is licensed under Apache 2.0 and built on top of an AI Catalog data model maintained through a working group under the Linux Foundation — a deliberate choice, and one that echoes a decision MCP itself made only a few months earlier: Anthropic donated MCP's own governance to a newly formed Agentic AI Foundation, a directed fund under the Linux Foundation, on December 9, 2025, with Anthropic, Block, and OpenAI as co-founders and Google, Microsoft, AWS, Cloudflare, and Bloomberg as supporting members. Discovery infrastructure that every vendor's agents are expected to rely on doesn't work if it's controlled by one of those vendors — ARD landing at the Linux Foundation from its first announcement, rather than migrating there later the way MCP did, suggests the industry learned that lesson the first time around.

ARD isn't a competitor to MCP — it's the layer MCP always assumed something else would eventually build: a way to find a server before you already know its address.

AGENTIC RESOURCE DISCOVERY

ai-catalog.json

example.com

ai-catalog.json

acme-corp.com

ai-catalog.json

another-domain

Registry

Crawls + indexes catalogs, searches by task intent

Agent

"what can handle a refund?"

“The problem is simple to state and hard to solve, especially in the enterprise, where the answer can't just be "find something that works." It has to be governed, with security and identity built in rather than bolted on.”

#09Four Problems Hiding Inside One Word: "Tool Discovery"

Everything covered so far — hard-coded tool lists, the Tool Search Tool, code execution with MCP, the MCP Registry, ARD's catalogs and registries — gets casually described as "tool discovery" in most conversations about agent tooling. That framing quietly flattens four genuinely distinct engineering problems into one word, and the flattening is exactly why so many teams solve one of them and assume they've solved all four.

An ARD registry answering "what exists" tells an agent nothing about whether a specific matching tool is the *right* one for this specific request among several plausible matches — that's a selection problem, and it's the one the tool-count-versus-accuracy research earlier in this piece was actually measuring. A Tool Search Tool correctly narrowing down to the right function tells an agent nothing about how to safely call it without leaking a two-hour transcript through its own context — that's an execution problem, and it's what code execution with MCP was built to address. And a tool an agent has both found and correctly selected might still be one it has no business calling at all — a refund tool with no spending limit attached, a database query tool with no row-level access control, a server whose catalog entry nobody has actually verified belongs to who it claims to. That's a trust problem, and none of the mechanisms discussed so far touch it directly. Treating "tool discovery" as one problem with one fix is how a team ends up with a fast, accurate agent that confidently calls a tool it should never have been allowed to reach. Broken apart, the one word people keep using actually names four separate questions:

"Tool discovery" isn't a single feature to add — it's an umbrella term for four different questions, and a system that answers only one of them will feel broken in a way that's hard to diagnose until you name the other three.

DISCOVER

"What tools exist, anywhere?"

SELECT

"Which one is right for this task, right now?"

EXECUTE

"How do I call it without breaking something?"

TRUST

"Should I even be allowed to use this?"

#010Discover: What Exists

Discovery, taken narrowly, is the question ARD is purpose-built to answer: given a task, what resources exist anywhere that could plausibly help — not just the handful a developer happened to hard-code, but everything a registry has indexed across every domain that's published a catalog. The two live reference implementations of the specification make the shape of this concrete rather than theoretical. GitHub's Agent Finder, shipped for Copilot on the same day the specification itself was announced, lets a developer describe a task in plain language; it searches an index of MCP servers, skills, canvases, and agents drawn from GitHub's own public catalog — or a private enterprise registry an organization points it at instead — and returns ranked matches Copilot can pull in on demand. Nothing gets auto-installed; a human still approves what actually gets activated, which matters more once the next two sections cover what "activated" is allowed to mean. Hugging Face's Discover Tool does the analogous thing for a different population of resources: semantic search across the thousands of ML applications, demos, and specialized research skills already living on the Hub, exposed through the same ARD-shaped discovery flow rather than a bespoke Hugging-Face-only search box.

What both implementations share is the part that actually matters for a production agent: discovery becomes a *query*, not a *deployment step*. Under the old model, adding a new capability to an agent meant a developer finding it, reading its docs, and shipping a config change — discovery and integration were the same act, performed once, by a human, ahead of time. Under ARD, discovery is something the agent itself can do at request time, against a catalog that keeps growing without the agent's own code ever needing to change. That shift is what makes the earlier 500-tools scenario tractable in the first place — an agent doesn't need all 500 tool definitions sitting in its context or its codebase; it needs a registry it can ask, and a way to act on what comes back, which is exactly where the next problem starts.

Discovery turns "does this agent know about that capability" from a question answered once, at deploy time, by a developer, into a question the agent can ask for itself, every time, against a catalog it never had to be told about in advance.

#011Select: What's Actually Appropriate Right Now

Finding twelve tools that plausibly match "handle a refund" is not the same as picking the one that's correct for *this* refund, for *this* customer, in *this* state, and that gap is exactly where the tool-count accuracy research from earlier in this piece lives. A registry search and a model's own tool-selection reasoning are doing genuinely different work: the registry narrows an unbounded space down to a shortlist worth considering at all; selection is the harder judgment call of picking correctly from that shortlist, under exactly the same pressures — attention dilution, tool collision, positional bias — that made accuracy collapse from 84–95% down to 0–20% as raw tool counts climbed in the benchmark cited earlier.

This is precisely the layer Anthropic's Tool Search Tool operates at, and it's worth being explicit that it solves a problem discovery alone cannot: even after ARD or an MCP Registry has narrowed the field to "these forty tools are plausibly relevant," an agent still needs *all forty definitions* usefully represented in its reasoning to choose correctly among them — and forty definitions is already enough to reintroduce meaningful accuracy loss if handed over naively. Semantic, on-demand retrieval — search first, expand only what's relevant, exactly as defer_loading does — is what keeps that narrowed shortlist from becoming its own smaller version of the same tool-collision problem. The two layers are genuinely complementary rather than redundant: a registry answers "what's out there," and a search-based selection mechanism answers "which of the things that are out there should actually go into this specific reasoning step" — skipping either one leaves the other doing a job it wasn't built for, whether that's a Tool Search Tool trying to search across every MCP server on the internet, or a global registry expected to also make the final judgment call a model's own reasoning is better positioned to make.

Mechanically, this is the same pattern that made RAG the default way to ground a model in a document collection too large to fit in context, just pointed at tool descriptions instead of paragraphs of text. Every tool's name, description, and parameter schema gets converted into an Embeddings vector ahead of time and stored in an index; a request like "handle a refund for this customer" gets embedded the same way at query time, and a nearest-neighbor search over that index returns the tools whose descriptions sit closest to the request in that vector space — not by keyword overlap, but by meaning, which is what lets "cancel this charge" and "issue a refund" retrieve the same tool even though they don't share a single word. That candidate list is still small enough — five or ten entries, not five hundred — for the model's own reasoning to make the final call reliably, which is exactly the accuracy effect Anthropic measured: the model wasn't getting smarter, it was simply being asked to choose from a list short enough that tool collision and positional bias stopped being a factor. The one thing this mechanism can't do on its own is rank by anything other than semantic similarity — a tool with a beautifully precise description but no track record of actually working reliably retrieves exactly as well as one that's been battle-tested in production, which is a gap the trust layer, not the selection layer, is what eventually has to close.

A registry narrows a search space from unbounded to plausible; selection narrows it one more time, from plausible to correct — and skipping straight from "everything" to "the one I'll actually call" is exactly what collapses accuracy at scale.

#012Execute: Using It Without Breaking Something

Correctly selecting a tool still leaves the question of what happens the moment it's actually called, and that's where code execution with MCP earns its place as a distinct pillar rather than a variation on selection. Traditional tool calling treats every call as a round trip through the model itself — the model emits a call, gets a result back into its own context, decides what to do with that result, emits the next call. That's fine for a single lookup. It becomes the meeting-transcript problem the moment a task genuinely needs to move real data between two systems: reading a document is one call, writing it somewhere else is another, and under the traditional model, the full content of that document has to pass through the model's context twice to get from one system to the other, whether or not the model needs to actually reason about the transcript's contents at all.

Executing inside a sandbox instead of inside the model's own reasoning loop changes what "using a tool safely" even means. It's not only a token optimization — though it clearly is one, at nearly 99% reduction on Anthropic's published example — it's also a containment boundary. Code the agent writes to move a document from Google Drive to Salesforce runs in an environment that can be resource-limited, time-limited, and denied network access to anything outside the two systems the task actually needs, independent of whatever the model itself decides to write. A tool call that used to mean "the model's next output directly triggers a side effect in a production system" becomes "the model's next output is code that runs somewhere isolated, with its own boundaries, before any side effect actually lands" — a meaningfully different security posture, and one that starts to blur into the trust question the next section covers directly.

The same execution environment quietly picks up a second job once it's in place: an agent can persist intermediate results to files inside the sandbox across separate turns of a task, rather than re-fetching or re-deriving the same data every time it's needed, and it can save the code it wrote to accomplish a task as a reusable function for the next time something similar comes up. Neither of those is possible when every action is a single, isolated call back into the model's own context — there's nowhere for a "the thing I built last time" to actually live. Give the agent a real execution environment instead of a call-and-response loop, and it starts accumulating the same kind of working infrastructure a human engineer would: helper functions, cached intermediate state, code worth reusing instead of rewriting from scratch on every single request.

Selecting the right tool answers "what should happen next" — execution is the separate, harder engineering question of "where does that actually run, and what can it touch while it does."

#013Trust: Should the Agent Even Be Allowed To

Everything up to this point assumes good faith on all sides: the catalog is accurate, the server is who it claims to be, and the agent calling a tool is authorized to call it for this particular user, on this particular data. None of the first three pillars actually verify that assumption, which is exactly why ARD builds trust in as a first-class part of discovery rather than an afterthought bolted on later. A catalog's domain-based identity — the fact that ai-catalog.json lives at a well-known path on a domain the publisher actually owns — is the specification's first trust layer: an agent resolving a catalog gets a cryptographically grounded answer to "is this really from who it claims to be," the same foundational guarantee HTTPS and domain ownership already provide for every other kind of web resource. The specification layers a trust-manifest schema and globally unique namespaced identifiers on top of that, giving a registry — and an agent consuming its results — a way to verify a resource's claimed identity before ever making a runtime connection to it, which is the third of the specification's four phases, deliberately sequenced *before* the actual connection happens rather than after.

Governance sits directly underneath that technical trust layer, and it's worth being explicit about why it matters here specifically. A discovery registry that every agent in an industry is expected to query is exactly the kind of infrastructure that becomes dangerous if one company quietly controls what gets ranked, surfaced, or hidden from it — which is precisely the concern MCP's own donation to the Agentic AI Foundation on December 9, 2025 was meant to preempt, months before ARD needed to answer the identical question for a much bigger surface area. That's also the level at which enterprise controls actually bite in production: GitHub's Agent Finder ships with governance settings that let an organization restrict which catalogs and registries its Copilot agents are even allowed to search in the first place, enforced in the same place the rest of Copilot's enterprise governance already lives — a company can point its agents exclusively at a private, internally-vetted registry rather than the open public one, which is a trust decision made once, centrally, rather than trusted to whatever an individual agent happens to discover on its own.

It's worth separating that catalog-level trust from a narrower, older question ARD deliberately leaves to the protocol underneath it: once an agent has verified a server is genuinely who it claims to be, is *this specific agent* actually authorized to call *this specific tool*, on behalf of *this specific user*, with *this specific data*? That's the question OAuth 2.1-based authorization, already part of the MCP specification for remote servers, was built to answer — a user grants a scoped, revocable token to a specific client, and the server checks that token on every call, independent of whatever a catalog said about the server's identity. The two checks aren't redundant: a catalog can correctly verify that payments.example.com really is payments.example.com and still have nothing to say about whether the particular agent asking to issue a refund actually has a customer's consent to do so. Trust, at real scale, ends up being both checks stacked — catalog-level identity answering "is this really the server it claims to be," and OAuth-style authorization answering "is this specific call actually permitted." Laid out side by side, the layers this section has covered stack like this:

Trust doesn't remove the need for the other three pillars — an agent still has to discover, select, and execute correctly even once it's confident a resource is legitimate. What it adds is the one question none of the other three are positioned to answer on their own: not "can this agent technically reach this tool," but "should it be allowed to," a distinction that matters enormously once discovery stops being a fixed list a developer personally vetted and becomes a live search across catalogs an agent has never seen before.

Discovery, selection, and execution all assume the tool in front of the agent is legitimate — trust is the pillar that actually checks, and it's the one most naive agent stacks skip entirely because nothing forces them to notice its absence until something goes wrong.

LayerWhat it answersWhere it lives in ARD
Domain identityIs this catalog really from who it claims?Cryptographic verification tied to domain ownership
Trust manifestWhat is this resource actually claiming to be able to do, and by whom?Metadata attached to each catalog entry
GovernanceWhich catalogs/registries is this agent even allowed to search?Enterprise-level controls (e.g., GitHub Agent Finder's managed settings)

#014Putting the Four Together: One Request, Start to Finish

Laid end to end, the four pillars aren't competing solutions to the same problem — they're sequential stages a single request actually passes through, each one handing a narrower, more verified result to the next. Go back to the opening scenario with that in mind: a support agent facing 500 tools across 50 MCP servers, asked to "find my last order," no longer has to reason over all 500 at once, or over whatever subset one developer happened to hard-code months ago. Discovery narrows 500 down to the handful of catalogs that even claim to expose order-related capability. Selection narrows that handful down to the one tool that actually matches this request. Trust confirms the server behind that tool is who it claims to be, and that this agent is actually authorized to call it for this customer. Execution runs the call somewhere contained, returning only the answer the user actually asked for — not the internal machinery it took to get there. Each stage is doing a small amount of work; the value is in none of them having to do all of it alone.

Skip any single stage and the pipeline doesn't just get slower — it fails in a specific, identifiable way. Discovery without selection means an agent finding forty plausible tools and then guessing among them with the same collapsing accuracy the earlier benchmark table showed. Selection without trust means an agent correctly picking the best-matching tool in a registry, with no verification that the registry entry is actually who it claims to be. Trust without execution discipline means an agent correctly verifying a legitimate tool and then still routing an entire two-hour transcript through its own context to call it. And execution without discovery is just the hard-coded tool list this piece started with — a real, working system, just one that only ever reaches the handful of tools someone remembered to wire up by hand. None of the four pillars is optional at real scale; each one is covering for a specific failure mode the others structurally can't catch. Traced through a single request, the whole pipeline looks like this:

A request that survives discovery, selection, trust, and execution intact isn't lucky — it's a request that passed through four separate checks, each one built to catch a different way the whole thing could have gone wrong.

"Find my last order"

DISCOVER

ARD registry search: which catalogs expose order-related capability?

SELECT

Tool Search Tool: of the matches, which specific tool fits this request?

TRUST

Verify catalog identity + trust manifest before connecting

EXECUTE

Call the tool inside a sandboxed environment; only the result the model actually needs returns to its context

Answer to the user

#015What This Doesn't Solve Yet

It would overstate things to treat any of this as finished. ARD is a June 2026 specification, still explicitly a draft, and its own documentation is candid that it deliberately doesn't specify a single global catalog or a single ranking algorithm — Marsman's "many discovery services" framing is a genuine design choice, not a gap, but it also means an agent's results depend entirely on which registries it happens to query and how each one chooses to rank matches, with no universal answer to "which registry is authoritative" for any given domain. Community discussion around the specification's early rollout has already surfaced the obvious next questions it doesn't yet answer on its own: how does an agent judge the *quality* of a tool a registry surfaces, beyond whatever the publisher claims about it, and what happens to pricing and access models once a tool can be discovered and invoked by agents its author never directly negotiated an integration with.

The Anthropic-side fixes have real edges too. The Tool Search Tool and code execution with MCP were both demonstrated on Anthropic's own models and benchmarks; the underlying pattern — search before load, execute in a sandbox rather than through the model's own context — is portable in principle, but the specific token and accuracy figures cited throughout this piece are Anthropic's own numbers, on Anthropic's own evaluations, not a vendor-neutral standard the way ARD's catalogs are meant to be. And the tool-selection-accuracy research showing a collapse to 0–20% at hundreds of tools is exactly the kind of number that varies by model, by benchmark design, and by how semantically similar the tools in a given catalog actually are to each other — a useful signal that the problem is real and severe, not a universal constant every agent will hit at exactly the same tool count.

None of that is a reason to wait. It's a reason to treat this the way any working engineer treats a new piece of infrastructure still finding its footing: adopt the pattern — discover before you hard-code, select before you load everything, verify before you trust, sandbox before you execute — while staying honest that the specific tools implementing each pillar today are early, and some of what gets built against them now will need to change as the specification and its reference implementations mature.

Every piece of this is a genuine improvement over what came before it — none of it is a finished, settled answer yet, and treating it as one is its own kind of risk.

#016What This Means If You're Building an Agent Today

None of the four pillars deserve equal urgency on day one, and reaching for all of them before an agent actually needs them is its own mistake — the same one hard-coding a tool list already avoided at small scale. If an agent genuinely only ever needs a dozen or two tools, all internal, all vetted by the same team that built the agent, the accuracy research in this piece simply doesn't apply yet: the naive hard-coded list from the second section of this piece is still the right answer, and reaching for a registry, a search tool, and a trust-manifest verification step at that scale adds real complexity for a problem that hasn't actually shown up.

The moment worth watching for is the one where a tool catalog stops being something one team fully understands — a second team starts contributing tools, an integration gets added that nobody who reads the agent's prompt actually remembers exists, or the total count crosses roughly the 50-to-100-tool range where the accuracy research starts showing real degradation. That's the trigger to introduce selection first — defer_loading, or an equivalent search-before-load pattern — since it's the cheapest of the four pillars to retrofit onto an existing hard-coded list without restructuring how tools get registered at all. Execution discipline is worth building in early specifically for any tool that moves meaningful data between systems or has a real side effect, even before the token savings matter, because sandboxed execution is a security boundary as much as an optimization, and it's considerably easier to design that boundary in from the start than to retrofit it once a tool is already calling production systems directly from inside a model's own reasoning loop. Discovery via an external registry — ARD, the MCP Registry, or an internal equivalent — earns its place once an agent's tool surface genuinely spans systems no single team controls end to end, which for most organizations is later than it feels like it should be. And trust verification isn't really optional at all once discovery goes external: the instant an agent can find and call a tool nobody on the team personally vetted, skipping verification isn't cutting a corner on a nice-to-have, it's removing the one check that was standing between "an agent that can reach a lot of tools" and "an agent that will call whatever a registry hands it."

Build for the tool count you actually have, not the one a roadmap says you'll have someday — but know exactly which of the four pillars to reach for first when that count actually starts climbing, because by the time it's a visible problem, retrofitting trust and execution discipline is considerably harder than retrofitting search.

#017The Real Shift

Every headline about AI agents this year has been about capability — more tools connected, more systems reachable, more autonomy granted. Nothing in this piece argues against that trajectory continuing. But capability was never actually the constraint the industry assumed it was; connecting a five-hundredth tool to an agent has been mechanically easy for a while now. What's been missing is the much less glamorous layer underneath — a way to find the right one of those five hundred without loading all of them, choose correctly once a shortlist exists, run the chosen one somewhere contained, and verify it was safe to reach for in the first place. That's not a capability problem. It's the same unglamorous discipline every mature piece of infrastructure eventually has to build: an index, a ranking step, a sandbox, and an identity check — the parts that don't show up in a product demo but are the entire difference between a demo and something a business can actually depend on.

It's also a familiar shape to anyone who's watched a fast-growing codebase go through the same transition. A small team's first API doesn't need a service registry, a schema validator, or a permissions layer — everyone who calls it also wrote it, and knows exactly what it does and doesn't do. None of that stays true once the API has fifty callers nobody who maintains it has ever met. Agent tooling is living through the identical shift, just compressed into months instead of years, because the tool count an agent accumulates grows far faster than the endpoint count any one API ever did. Discover, select, execute, trust isn't a new idea invented for AI agents — it's the same maturity curve every piece of shared infrastructure eventually climbs, arriving now because agents finally have enough tools connected to need it.

The interesting problem was never how many tools an agent could reach — it was always going to be whether the layer underneath could tell it, correctly and safely, which one to reach for.

The future of agents may not be about giving them access to everything. It may be about teaching them what they should access — and why.

Found this useful? Share it

Have a technical response or architectural perspective to share with the engineering desk?

Submit Engineering Feedback