AI Agent Evaluation: How to Test Accuracy, Reliability and Performance
How to evaluate AI agents: building evaluation datasets, task success, tool-call accuracy, groundedness, policy compliance, latency, cost, LLM-as-judge, regression testing and production evaluation.
Quick answer
Evaluate AI agents on whole tasks, not single answers. Build a dataset of real cases with expected outcomes, run the agent against it, and score the final result (task success, correct end state), the trajectory (tool choice, argument validity, step count, policy compliance), output quality (groundedness, tone) and operations (latency, cost per task). Use deterministic checks wherever possible, LLM-as-judge with human calibration where needed, gate every release on regression results and keep scoring sampled production runs.
Where This Fits
Evaluation needs traces from agent observability, informs the thresholds in human-in-the-loop design and tests the controls in guardrails. For retrieval-specific metrics, see RAG.
Evaluation of individual models and prompts, rather than whole agents, is covered in AI model evaluation.
What to Measure
| Level | Metric | How to score |
|---|---|---|
| Outcome | Task success | Compare end state or output with expected result |
| Outcome | Escalation correctness | Did it hand off when it should, and only then? |
| Trajectory | Tool-call accuracy | Right tool, valid and correct arguments |
| Trajectory | Efficiency | Steps, loops, redundant calls |
| Quality | Groundedness | Claims supported by sources or tool results |
| Safety | Policy compliance | No forbidden actions or disclosures |
| Operations | Latency and cost per task | Measured per run |
Building an Evaluation Dataset
Start from real cases, anonymized where needed. Include typical cases, known edge cases, cases where the right answer is to escalate, and adversarial inputs such as prompt injection attempts. For each case record the input, relevant context or system state, the expected outcome and, where useful, expected tool calls. Label who agreed the expected outcome. Keep a held-out set you do not tune against.
- Typical cases weighted by real volume
- Edge cases from support tickets and incident reports
- Should-escalate cases
- Adversarial and malformed inputs
- Cases that exercise every tool and policy
- A held-out set for honest comparison
Scoring Methods
Prefer deterministic checks: did the order status become 'cancelled'? Does the output match the schema? Was the refund under the limit? Use reference comparison for extracted fields. Use LLM-as-judge for qualities such as helpfulness or groundedness, with a written rubric, and calibrate it by comparing its scores with human ratings on a sample. Use human review for high-stakes or ambiguous cases.
Evaluating Tool Use and Trajectories
For agents, the path matters. An agent that reaches the right answer by calling a forbidden tool has failed. Score each step: was the tool appropriate, were arguments valid, was the call necessary? Compare against expected calls where the path is fixed, and against state checks where several paths are acceptable. Flag loops and excessive steps, which drive cost and latency.
Shipping agent changes without knowing what they break?
ZSpace Labs sets up evaluation datasets, scoring and release gates so every prompt, model or tool change is tested against real cases.
Regression Testing and Release Gates
Run the evaluation set on every change to prompts, tools, retrieval or model version, and compare against the current production version. Define gates: overall success must not drop, safety failures must be zero, cost and latency must stay within budgets. Track results by case type, because averages hide regressions in small but important categories.
Rollout strategies after the gate, such as shadow tests and canaries, are covered in AI application release management.
Production Evaluation
Offline sets never capture everything. In production, score a sample of live runs automatically, collect user and reviewer feedback, watch outcome signals (corrections, escalations, complaints) and review failures weekly. Turn every confirmed failure into a new test case. Store traces with privacy controls, because they contain customer data.
Tools for Agent Evaluation
Options include evaluation features in model providers' platforms, open-source frameworks, observability platforms with evaluation built in, and simple custom harnesses that replay cases and score results. Choose tools that store datasets with versions, run on every change, and link scores to traces so failures can be debugged.
Advantages and Limitations of Evaluation Approaches
| Method | Strengths | Limitations |
|---|---|---|
| Deterministic checks | Exact, cheap, repeatable | Only for checkable outcomes |
| Reference comparison | Clear for extraction tasks | Needs labelled answers |
| LLM-as-judge | Scales to open-ended quality | Bias, needs calibration |
| Human review | Most trusted | Slow and costly |
| Production signals | Real behaviour | Lagging and noisy |
How to Set Up Evaluation Step by Step
- 1. Agree success criteria with the process owner
- 2. Collect and label 50 to 200 real cases
- 3. Write deterministic checks for end states and schemas
- 4. Add rubric-based judging for open-ended quality, calibrated against humans
- 5. Run a baseline and record results by case type
- 6. Add the run to CI as a release gate
- 7. Sample production runs and feed failures back into the set
Anatomy of an Evaluation Case
A good evaluation case records enough to replay the task and judge it objectively. Store cases as versioned data alongside the code they test.
{
"id": "refund-017",
"input": "I was charged twice for order ORD-104233, please fix it",
"setup": { "order": "ORD-104233", "payments": ["pi_a", "pi_b"], "delivered_days_ago": 3 },
"expected": {
"outcome": "refund_issued",
"refund_amount": 59.90,
"tools_required": ["get_order", "list_payments", "refund_payment"],
"tools_forbidden": ["cancel_order"],
"must_escalate": false
},
"checks": ["end_state_refund_count == 1", "reply_mentions_refund_timeline"],
"labelled_by": "support-lead",
"tags": ["refund", "duplicate_charge"]
}Evaluating Retrieval Inside Agents
Many agents depend on retrieval for policies or records. When an agent gives a wrong answer, check whether it retrieved the right sources before blaming the model. Add retrieval checks to evaluation cases (which sources should appear?) and measure faithfulness of the final answer to what was retrieved. The RAG-specific metrics are covered in the RAG guide.
Who Owns Evaluation
Evaluation works when ownership is clear. The process owner agrees what correct looks like and approves labels; engineers maintain the harness, CI gates and production sampling; reviewers or QA staff label new cases from production failures. Review results together after each release and monthly for trends.
Worked Example
An illustrative scenario, not a client case: a support agent passes informal testing, but an evaluation set of 250 real tickets shows it issues refunds on orders outside the return window in a small share of cases. The fix is a policy check inside the refund tool, not a prompt change. The case is added to the regression set, and the release gate now requires zero policy violations.
Common Mistakes
- Testing with a handful of invented examples
- Scoring only the final answer, not the actions
- Trusting an uncalibrated LLM judge
- Averages that hide failures in important case types
- No evaluation when the model provider updates a model
- Not adding production failures to the test set
Want evidence that your agent is ready for production?
Talk to ZSpace Labs about AI evaluation and agent development.
Conclusion
Agent evaluation is how you earn the right to automate. Build datasets from real cases, score outcomes and trajectories, gate releases and keep learning from production. Related: observability, guardrails and agent development.
Common questions
Measuring whether an agent completes tasks correctly, safely and efficiently, using test cases with expected outcomes, scoring of each step and the final result, and comparison across versions.