Skip to content
AI & Automation7 min read

AI Agent Tool Design: How to Build Tools an Agent Can Use Reliably

How to design tools an AI agent can use reliably: granularity, names and descriptions, input schemas, outputs, errors, safe writes and evaluation.

01

Quick answer

An agent is only as good as the actions it can take. Design tools around tasks, not raw endpoints; keep the set small and distinct; give each tool a clear name, a description that says when to use it and a strict input schema; return compact, meaningful results and errors that explain how to recover; separate reads from writes; make writes idempotent; enforce permissions and limits in code; and require confirmation for consequential actions. Then evaluate the tools on real tasks and refine them from the transcripts.

02

Where tools fit in an agent

Every major model API supports tool use (also called function calling): you describe functions, the model decides when to call one and with which arguments, your application runs it and returns the result, and the loop continues until the task is done. The Model Context Protocol standardizes the same idea across applications, so a tool written as an MCP server can be used by many agents.

The model never touches your systems directly. Your code does, which means tool design is where usefulness and safety are decided. For the overall build, see our AI agent development guide; this article focuses on the tools themselves.

03

Choose the right granularity

The most common mistake is wrapping every API endpoint as a tool. An agent then has to discover multi-step workflows, juggle IDs and read large payloads, which wastes context and creates errors. Anthropic's engineering guidance on writing tools for agents makes the same point: build a few thoughtful tools for high-impact workflows rather than mirroring an API.

Endpoint-shaped toolTask-shaped toolWhy it is better
list_customers + list_orders + get_orderfind_customer_orders(customer_email, status)One call, no ID juggling
update_record(table, id, fields)update_shipping_address(order_id, address)Narrow, validatable, auditable
run_sql(query)get_sales_summary(period, region)No arbitrary queries; predictable cost
send_email(to, subject, body)send_order_update(order_id, template)Recipients and content constrained

Pro tip

Write down the ten tasks the agent must handle, then design the smallest set of tools that covers them. Add tools only when a real task needs one.

04

Names and descriptions are prompts

The model chooses tools from their names and descriptions, so treat them as carefully as any prompt. Use clear verbs and nouns, consistent prefixes for related tools (`orders_search`, `orders_get`), and descriptions that say what the tool does, when to use it, when not to, and what it returns. Spell out units, formats and constraints.

A tool definition with a useful description
{
  "name": "orders_find_by_customer",
  "description": "Find a customer's orders by email. Use when the user asks about their orders, deliveries or returns. Returns up to 10 most recent orders with status and delivery date. Does not return payment details. For a single known order number, use orders_get instead.",
  "input_schema": {
    "type": "object",
    "properties": {
      "customer_email": { "type": "string", "format": "email" },
      "status": { "type": "string", "enum": ["any", "open", "shipped", "delivered", "returned"], "default": "any" }
    },
    "required": ["customer_email"],
    "additionalProperties": false
  }
}
05

Make inputs hard to get wrong

Use strict schemas: enums instead of free text, explicit formats for dates and emails, required fields marked, and no unexpected properties. Most model APIs offer a strict or structured mode that constrains arguments to the schema; turn it on where available. Then validate again on the server anyway, because the schema guides the model but your code is the enforcement point.

Prefer identifiers the model can obtain naturally (an email the user gave, an order number shown in an earlier result) over internal IDs it would have to guess.

06

Return what the next step needs

Tool results go back into the model's context, where every token costs money and attention. Return the fields that matter for the next decision, use human-readable names alongside IDs, paginate or truncate long lists and say that you did, and offer a concise and a detailed mode if both are needed.

Errors deserve the same care. "400 Bad Request" teaches the agent nothing. "No customer found for that email. Ask the user to confirm the email address or provide an order number" lets it recover.

Weak resultBetter result
Full JSON of 200 orders with 60 fields each10 most recent orders: number, date, status, total, delivery date; note that more exist
{"error": "ERR_42"}"Address update not allowed: order already shipped. Offer the user a return or redirect request instead."
Internal status code 7"Status: awaiting payment confirmation"
07

Design writes for safety and retries

Write tools change the world, so they need more than a good description.

  • Separate reads from writes so read-only agents can be given only read tools
  • Make writes idempotent with an idempotency key, so a retried call does not create two refunds
  • Enforce authorization in code using the identity of the user the agent acts for, never a flag in the prompt
  • Limit blast radius: maximum amounts, quantities and affected records per call
  • Require confirmation for destructive, financial or external-communication actions, ideally by returning a preview the user approves
  • Log every call with arguments, result, identity and the run it belonged to

Worth noting

MCP lets servers annotate tools with hints such as read-only or destructive. They help clients decide when to ask for confirmation, but they are hints. Enforcement still belongs in your server.

08

Confirmation patterns that work

Asking "Are you sure?" before every action trains people to click yes. Reserve confirmation for actions that are costly, irreversible or visible to others, and make it informative.

PatternHow it worksUse for
Preview then commitTool returns a summary and a token; a second call with the token executesRefunds, bulk updates, sending messages
Draft for humanAgent creates a draft in your system; a person sends or applies itCustomer emails, contracts, price changes
ThresholdsAuto-execute below a limit, require approval above itDiscounts, credits, purchase orders
Undo windowExecute with a short reversal periodLow-risk, reversible changes

Designing agents that take real actions?

ZSpace Labs builds task-shaped tools, MCP servers and approval flows around your existing systems. See AI automation services.

Start a Project
09

Evaluate tools with real tasks

Tools are hard to judge by reading them. Build a set of realistic requests (including ambiguous and adversarial ones), run the agent, and check whether it chose the right tools with the right arguments, how many calls it needed and where it failed. Read the transcripts: confusion between two tools usually means their descriptions overlap; repeated retries usually mean errors are unhelpful; long runs usually mean results are too verbose.

Repeat the evaluation whenever you change a tool, its description or the model. Our AI agent evaluation guide covers building the test set and metrics.

10

Common mistakes

  • Mirroring every API endpoint as a separate tool
  • Overlapping tools with similar names and vague descriptions
  • Free-text parameters where an enum or format would do
  • Returning entire records or unpaginated lists
  • Opaque error codes the agent cannot act on
  • Write tools without idempotency, limits or confirmation
  • Relying on the prompt to enforce permissions
  • Never evaluating tools against real tasks
11

Conclusion

Good agents are built from good tools: few, task-shaped, clearly described, strictly typed, economical in what they return and safe when they write. Treat tool design as product design for a very literal user, evaluate it with real tasks and enforce every rule in code. For the security side, read AI tool security; to package tools for many agents, see how to build an MCP server.

When an agent has access to many tools, see AI agent tool selection; to make tool arguments and results dependable, see structured outputs.

FAQ

Common questions.

A function the model can ask your application to run, described by a name, a natural-language description and an input schema. The model chooses a tool and arguments; your code executes it and returns the result.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.