Evaluating a production IT support agent without an answer key
Adapting recent trajectory-evaluation research to a real tool-calling agent, then using it to compare model configurations on quality, cost, and latency.
Context. This study comes from hands-on work evaluating a production enterprise IT support agent, published with the organization’s permission. Organization, product, data, and identifying details are withheld; all figures are rounded or expressed relatively.
The problem
The agent answers employees’ IT questions by calling tools: a ticketing system, an employee directory, and a knowledge base. The team needed to know how well it behaved on real traffic, and whether cheaper or faster model configurations would hold up.
Two obstacles made this hard:
- No answer key. Real support conversations have no ground-truth reply or reference trajectory. Standard “compare to the expected answer” metrics do not apply.
- Final answers hide the failures. An agent can reach a plausible reply through fabricated arguments, redundant tool calls, or ignored errors. Judging only the last message misses most of that.
Research foundation
No off-the-shelf framework covered step-level, reference-free evaluation of this kind of agent, so the method was assembled from recent papers:
| Paper | Idea used |
|---|---|
| TRACE — Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents (ICML 2026) | An evidence bank built from prior tool calls and results; step-level efficiency, hallucination, and adaptivity metrics |
| Agent GPA — What Is Your Agent’s GPA? (arXiv:2510.08847) | Judging goal–action alignment directly, without needing an exposed plan |
| Agent-as-a-Judge survey (arXiv:2601.05111) | Tool-augmented verification of factual claims (roadmap item) |
What had to change for a real agent
Published methods assume research-style trajectories. Production agents differ in ways that break them:
- No visible reasoning. TRACE judges an explicit “thought” against the evidence. Modern production models hide their reasoning, so the prompts were rewritten to judge the action itself: the tool name and arguments, or the claims in the final reply.
- Parallel tool calls. A single step can issue several calls at once. The whole bundle is judged as one step, including redundancy within the bundle.
- Noise in the logs. Exported traces repeat the full tool catalog on every call, and the agent is required to call bookkeeping tools. Tool definitions are stripped, and bookkeeping-only steps are skipped structurally, with no judge call and no cost.
- Adaptivity only when it applies. Recovery is scored only when the preceding tool result is actually a failure; otherwise the metric is N/A instead of a free pass.
The four metrics
- Groundedness: are tool arguments and reply claims supported by the user’s input and the evidence gathered so far?
- Efficiency: does the step avoid redundant or repeated calls?
- Adaptivity: after a tool failure, does the agent adjust sensibly?
- Goal alignment: does the step move toward what the user actually asked?
Making the judges trustworthy
A single LLM judge is noisy, so every step was scored by three judges from three different model providers, at temperature 0:
- Per step: 2-of-3 majority vote per metric.
- Per turn: a turn counts as negative on a metric if any of its steps is.
- Agreement tracking: three-way agreement is reported per metric, so readers know which numbers to trust.
What agreement revealed:
| Metric | Three-way agreement | Reading |
|---|---|---|
| Groundedness | ~77–97% | Reliable |
| Efficiency | ~89–95% | Reliable |
| Adaptivity | ~99% | Reliable, but rarely triggered |
| Goal alignment | ~70–76% on replays | Real disagreement; split cases go to human review |
Judges are not consistently strict or lenient. One judge was the most lenient on one dataset and the strictest on another. No single judge can be treated as calibrated, which is the strongest argument for an ensemble.
Using it to compare model configurations
The same set of real conversation turns was replayed against the agent under three backends, each judged the same way:
| Configuration | Quality | Cost | Latency |
|---|---|---|---|
| Model router (mix of large / mid / small tiers) | Best: about a quarter fewer turns with an ungrounded or misaligned step | ≈ fixed mid tier | Slowest tail |
| Fixed small tier | Lower | ~9× cheaper | Middle |
| Fixed mid tier | Lower (≈ small tier) | ≈ router | Fastest, most consistent |
No configuration won on all three. The right choice depends on which constraint matters most. That is a business decision, and the evaluation makes it an informed one.
Two further findings:
- Routing breaks bad streaks. The router switched models on about half of turn-to-turn transitions. Within multi-turn conversations, its rate of negative turns was about 30% lower than the fixed backends’, which failed the same hard conversations repeatedly.
- Small models are not automatically fast. The small tier reasoned longer per step than the mid tier, so it was cheaper but not quicker.
Lessons for anyone doing this
- Replay setup changes the agent’s behavior. Embedding prior tool results in the replayed history let the agent skip tool calls, cutting steps per turn by more than half. Replay results must be compared with each other, not naively with production.
- Check that a metric measures the agent, not the environment. In a simulated-user pass, a “goal fulfilled” metric looked terrible because a test-environment service returned empty data. Per-step goal alignment showed the agent was behaving reasonably.
- Reasoning judges need large output budgets. Tight token limits made reasoning models spend the budget thinking and return empty verdicts.
- Some inputs break some judges. One long transcript made one judge produce degenerate output every time, across all runs. An ensemble absorbs this, and a single-judge pipeline would silently lose the data point.
- Checkpoint everything. Long judge runs hit gateway resets. Per-step error tolerance and resumable checkpoints turn a crash into a refill.
Limitations and next steps
- Final-answer correctness is not verified against live systems. The next step is tool-augmented verification: the judge re-queries the same read-only tools to check the agent’s factual claims.
- Adaptivity needs failure-rich traffic; clean test environments rarely trigger it.
- Goal-alignment disagreements need a small human-labeled set to calibrate the judges.
What a client gets
A reference-free evaluation of their own agent traffic: step-level quality metrics, judge-agreement reporting, and a side-by-side comparison of model configurations on quality, cost, and latency, ending in a recommendation.