Context. This study comes from hands-on work evaluating a production enterprise IT support agent, published with the organization’s permission. Organization, product, data, and identifying details are withheld; all figures are rounded or expressed relatively.

The problem

The agent answers employees’ IT questions by calling tools: a ticketing system, an employee directory, and a knowledge base. The team needed to know how well it behaved on real traffic, and whether cheaper or faster model configurations would hold up.

Two obstacles made this hard:

  • No answer key. Real support conversations have no ground-truth reply or reference trajectory. Standard “compare to the expected answer” metrics do not apply.
  • Final answers hide the failures. An agent can reach a plausible reply through fabricated arguments, redundant tool calls, or ignored errors. Judging only the last message misses most of that.

Research foundation

No off-the-shelf framework covered step-level, reference-free evaluation of this kind of agent, so the method was assembled from recent papers:

Paper Idea used
TRACE — Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents (ICML 2026) An evidence bank built from prior tool calls and results; step-level efficiency, hallucination, and adaptivity metrics
Agent GPA — What Is Your Agent’s GPA? (arXiv:2510.08847) Judging goal–action alignment directly, without needing an exposed plan
Agent-as-a-Judge survey (arXiv:2601.05111) Tool-augmented verification of factual claims (roadmap item)

What had to change for a real agent

Published methods assume research-style trajectories. Production agents differ in ways that break them:

  1. No visible reasoning. TRACE judges an explicit “thought” against the evidence. Modern production models hide their reasoning, so the prompts were rewritten to judge the action itself: the tool name and arguments, or the claims in the final reply.
  2. Parallel tool calls. A single step can issue several calls at once. The whole bundle is judged as one step, including redundancy within the bundle.
  3. Noise in the logs. Exported traces repeat the full tool catalog on every call, and the agent is required to call bookkeeping tools. Tool definitions are stripped, and bookkeeping-only steps are skipped structurally, with no judge call and no cost.
  4. Adaptivity only when it applies. Recovery is scored only when the preceding tool result is actually a failure; otherwise the metric is N/A instead of a free pass.

The four metrics

  • Groundedness: are tool arguments and reply claims supported by the user’s input and the evidence gathered so far?
  • Efficiency: does the step avoid redundant or repeated calls?
  • Adaptivity: after a tool failure, does the agent adjust sensibly?
  • Goal alignment: does the step move toward what the user actually asked?

Making the judges trustworthy

A single LLM judge is noisy, so every step was scored by three judges from three different model providers, at temperature 0:

  • Per step: 2-of-3 majority vote per metric.
  • Per turn: a turn counts as negative on a metric if any of its steps is.
  • Agreement tracking: three-way agreement is reported per metric, so readers know which numbers to trust.

What agreement revealed:

Metric Three-way agreement Reading
Groundedness ~77–97% Reliable
Efficiency ~89–95% Reliable
Adaptivity ~99% Reliable, but rarely triggered
Goal alignment ~70–76% on replays Real disagreement; split cases go to human review

Judges are not consistently strict or lenient. One judge was the most lenient on one dataset and the strictest on another. No single judge can be treated as calibrated, which is the strongest argument for an ensemble.

Using it to compare model configurations

The same set of real conversation turns was replayed against the agent under three backends, each judged the same way:

Configuration Quality Cost Latency
Model router (mix of large / mid / small tiers) Best: about a quarter fewer turns with an ungrounded or misaligned step ≈ fixed mid tier Slowest tail
Fixed small tier Lower ~9× cheaper Middle
Fixed mid tier Lower (≈ small tier) ≈ router Fastest, most consistent

No configuration won on all three. The right choice depends on which constraint matters most. That is a business decision, and the evaluation makes it an informed one.

Two further findings:

  • Routing breaks bad streaks. The router switched models on about half of turn-to-turn transitions. Within multi-turn conversations, its rate of negative turns was about 30% lower than the fixed backends’, which failed the same hard conversations repeatedly.
  • Small models are not automatically fast. The small tier reasoned longer per step than the mid tier, so it was cheaper but not quicker.

Lessons for anyone doing this

  • Replay setup changes the agent’s behavior. Embedding prior tool results in the replayed history let the agent skip tool calls, cutting steps per turn by more than half. Replay results must be compared with each other, not naively with production.
  • Check that a metric measures the agent, not the environment. In a simulated-user pass, a “goal fulfilled” metric looked terrible because a test-environment service returned empty data. Per-step goal alignment showed the agent was behaving reasonably.
  • Reasoning judges need large output budgets. Tight token limits made reasoning models spend the budget thinking and return empty verdicts.
  • Some inputs break some judges. One long transcript made one judge produce degenerate output every time, across all runs. An ensemble absorbs this, and a single-judge pipeline would silently lose the data point.
  • Checkpoint everything. Long judge runs hit gateway resets. Per-step error tolerance and resumable checkpoints turn a crash into a refill.

Limitations and next steps

  • Final-answer correctness is not verified against live systems. The next step is tool-augmented verification: the judge re-queries the same read-only tools to check the agent’s factual claims.
  • Adaptivity needs failure-rich traffic; clean test environments rarely trigger it.
  • Goal-alignment disagreements need a small human-labeled set to calibrate the judges.

What a client gets

A reference-free evaluation of their own agent traffic: step-level quality metrics, judge-agreement reporting, and a side-by-side comparison of model configurations on quality, cost, and latency, ending in a recommendation.

Want this for your agent?

Tell us what it does and what decision you're facing. We'll reply with how we would approach it.

Get in touch