AI Agent Tool Selection: Why Too Many Tools Make Agents Worse
How AI agents choose which tool to call, why large toolsets reduce accuracy and cost more, and how routing, tool search and scoped toolsets fix it.
Quick answer
An agent chooses a tool by matching the task against the names, descriptions and schemas of the tools in its context. The more tools it sees, especially overlapping ones, the more often it picks the wrong one, and the more each request costs. Keep toolsets small and distinct per task, filter tools by permission before the model sees them, route deterministic choices in code, and use tool search or deferred loading when a large library is unavoidable. Then measure tool-selection accuracy on real tasks and fix descriptions that cause confusion.
How tool selection actually works
The model never inspects your code. On each step it sees a list of tool definitions (name, description, input schema) alongside the instructions, conversation and previous tool results, and produces either text or a structured request to call one of those tools. Everything that influences its choice is in that context: how clearly tools are described, how similar they are to each other, how many there are and what the task looks like. That is why tool selection is as much a context problem as a model problem; see context engineering.
Why more tools make agents worse
Connecting an agent to several MCP servers can add dozens or hundreds of tools at once. Anthropic's engineering team has described setups where tool definitions alone consumed tens of thousands of tokens before the agent read the user's request, and reported that letting the model search for relevant tools instead of loading all of them improved accuracy on its tool-use evaluations. OpenAI's function-calling guidance similarly suggests keeping the functions available at the start of a turn small, with tool search to defer large or rarely used parts of the tool surface.
| Problem | Effect |
|---|---|
| Overlapping tools (search_orders, find_order, lookup_order_status) | Wrong tool chosen; inconsistent behaviour |
| Large definitions loaded every request | Higher cost and latency; less room for task context |
| Tools irrelevant to the task | Distraction; occasional bizarre calls |
| Tools the user may not use | Security risk and refused calls that derail the run |
Key takeaway
Adding a tool is not free. Each one should earn its place by covering a real task that no other tool covers.
Five ways to improve tool selection
| Technique | How it works | When to use |
|---|---|---|
| Fewer, task-shaped tools | Merge endpoint-level tools into task-level ones (see AI agent tool design) | Always |
| Permission filtering | Only expose tools the current user, tenant and task may use | Always |
| Deterministic routing | Code picks the toolset or workflow from request type before the model runs | When categories are clear |
| Tool search / deferred loading | Agent searches a catalogue and loads only relevant definitions | Large libraries, many MCP servers |
| Specialist sub-agents | Each sub-agent has a small toolset for one domain | Several distinct domains in one product |
Factors beyond relevance
The best tool for a step is not only the most relevant one. When several tools could work, encode preferences explicitly rather than hoping the model infers them:
- Cost: prefer cached or internal lookups over paid external APIs
- Latency: prefer fast reads for interactive replies
- Reliability: avoid tools with known instability unless necessary; route around them when a circuit breaker is open
- Risk: prefer read-only tools; require confirmation for write tools
- Confidence: if the agent cannot decide between tools, ask a clarifying question instead of guessing
Measure and fix selection errors
Add tool-selection checks to your evaluation set: for each test case, record which tool should be called with which arguments, then measure how often the agent gets it right (see AI agent evaluation). Read transcripts of wrong choices. Most fixes are descriptive: rename tools, state when not to use each one, merge near-duplicates, or remove tools nobody needs. Re-run the evaluation after every toolset change, including when an MCP server you depend on updates its tools; see MCP governance.
Agents calling the wrong tools?
ZSpace Labs audits agent toolsets, restructures tools and routing, and adds tool-selection evaluation so changes are measurable. See AI automation services.
Common mistakes
- Connecting every available MCP server "in case it is useful"
- Several tools with nearly identical names and descriptions
- Exposing admin tools to every user's agent session
- Letting the model choose when a simple rule would do
- Never testing tool choice separately from final answers
Conclusion
Tool selection improves when the agent sees fewer, clearer, permitted tools. Design task-shaped tools, filter by permission, route in code where you can, use tool search for large libraries and measure selection accuracy. For making the calls themselves dependable, see structured outputs and AI tool security.
Common questions.
The model reads the names, descriptions and input schemas of the tools available in its context, compares them with the task and conversation so far, and outputs a call to the tool that best fits. It does not see your code, so descriptions and the size of the toolset largely determine how well it chooses.