Agent evaluation on every trace, without the LLM-judge bill
Judging every production trace with an LLM is too expensive, and sampling misses the failures that matter. Triage checks arranged as a decision tree give you full coverage at a fraction of the cost.
Most teams check a small sample of production traces with an LLM judge and hope the rest look the same. They usually don't. The failures that hurt, a refund issued outside policy or a tool call that quietly returned nothing, sit in the traces nobody looked at.
Judging every trace with a frontier model fixes coverage and breaks the budget. This article shows a third option: triage checks arranged as a decision tree, where cheap checks settle most traces and the LLM judge only sees the hard ones.
You will get the four levels of the tree, what belongs at each level, rules for ordering checks, and a worked cost comparison you can adapt to your own volume. The method works with any stack. At the end we show how Trodo applies it.
Why judging every trace with an LLM gets expensive
Agent evaluation in production has a volume problem. A single agent trace can hold a dozen LLM calls, several tool calls and thousands of tokens of context. A judge has to read most of that to decide anything.
Take an illustrative example. Say your support agent handles 50,000 traces a day. Each judge call reads about 6,000 input tokens (the trace plus a rubric) and writes 300 tokens of reasoning and verdict. At an example price of $3 per million input tokens and $15 per million output tokens, one judgment costs about $0.0225.
- One criterion per trace: about $1,125 a day, or roughly $34,000 a month.
- Three criteria (policy, grounding, tone), each judged separately: about $3,375 a day, or roughly $101,000 a month.
- Latency: each judgment adds a few seconds, so results lag behind traffic and queue up at peaks.
That bill can exceed what the agent itself costs to serve. So teams cut coverage.
Why a sample is not enough
The usual fix is to judge a 5% sample. That keeps the bill down, but agent failures are rarely spread evenly. The dangerous ones are rare and specific.
Say the refund agent issues an out-of-policy refund in 1 of every 2,000 traces. At 50,000 traces a day that is 25 bad refunds. A 5% sample sees about one of them per day, and whether that one gets flagged depends on the judge being right on that trace. You find the pattern days later, from finance, not from your checks. We go deeper on this trade-off in trace sampling vs full coverage.
The goal is 100% coverage at a cost closer to the sample. That needs a different shape of evaluation, not just a cheaper judge.
Triage checks: agent evaluation as a decision tree
Triage works the way an emergency room does. Not every patient needs a specialist. A quick check settles most cases, and only the unclear ones move up.
Applied to traces, each level of the tree does one of three things for a given criterion: pass the trace, fail it, or pass it up because it cannot tell. Levels are ordered from cheapest to most expensive.

- Code checks: deterministic rules over the trace structure. Microseconds, effectively free.
- Semantic checks: embeddings, similarity and small NLI models. Milliseconds, fractions of a cent.
- A small trained model: a classifier trained on your labeled traces, returning a calibrated pass/fail probability.
- An LLM judge: a full model with a rubric, for the traces nothing else could settle.
Level 1: code checks
Anything you can decide by reading fields belongs here. Code checks are exact, and they break loudly when the trace shape changes, which is useful too.
- A tool span returned an error status or an empty result.
- The final output fails schema validation.
- An action tool was called before the lookup it depends on. For the refund agent: issue_refund appears before check_refund_policy, or without it.
- Numeric limits: refund amount above the order total, more than N tool calls, the same call repeated in a loop.
- Latency or token budgets exceeded.
Here is a policy check for the refund agent, written as plain code:
Notice the check has three outcomes. None means this check has no opinion, and the trace moves on to the next level. A code check should never guess.
Level 2: semantic checks
Some questions are about meaning, not structure, but still don't need a large model. Semantic checks use embeddings and small models that answer narrow questions quickly.
- Grounding: does each claim in the answer have a close match in the retrieved documents or tool outputs? Low overlap on a factual claim is a strong signal.
- Intent match: is the answer about what the user asked? Compare the embeddings of the question and the reply.
- Known bad patterns: similarity to a library of past failing answers, such as promising a refund the agent cannot give.
- Entailment: a small NLI model checks whether the tool output supports the sentence the agent wrote.
Semantic checks settle clear cases in both directions. An answer that restates the policy tool's output closely passes. An answer with no support in any tool output fails. The middle band moves up.
Level 3: a small trained model
Once you have a few hundred labeled traces for a criterion, you can train a classifier for it. It is far cheaper than a judge and, on the narrow task it was trained for, often more consistent.
The key property is calibration. When the model says 0.95 probability of pass, about 95 in 100 of those traces should really pass. With calibrated scores you can set two thresholds: above the top one, pass; below the bottom one, fail; in between, escalate. You pick the thresholds from the error rate you can accept.
Level 4: the LLM judge
The judge gets what is left: traces where the rules had no opinion, the semantic signals were mixed and the classifier was unsure. These are the genuinely ambiguous cases, and they are where a judge's reasoning is worth paying for. Give it a narrow rubric and the evidence the earlier levels collected. For how judges fail and how to calibrate one, see LLM as a judge: accuracy, bias and cost.
How to order checks
The order decides the cost. A few rules cover most cases:
- Cheapest first. A check that costs nothing should always go before one that costs a cent.
- Most decisive first, within a level. If one code check settles 60% of traces for a criterion, put it ahead of one that settles 5%.
- Let hard failures stop the tree. If a tool call errored and the agent answered anyway, you already have a failure. You don't need a judge to confirm it.
- Escalate on uncertainty, not by default. A trace should reach the next level only because the current level could not decide.
- Scope checks to the traces they apply to. A refund policy check has nothing to say about a password reset. Skip it instead of scoring it.
Worked example: three refund traces
A customer asks for a refund on an order delivered 41 days ago. Follow the trace through the tree for the policy criterion:
- Code: the agent called check_refund_policy before issue_refund. The order is correct, so no failure yet.
- Code: the policy tool returned eligible: false, and issue_refund was still called. Fail. The tree stops here.
A second trace: same question, but the agent refused politely and offered store credit. Code checks find no refund, so the policy check has nothing to fail. The tone criterion goes to the semantic level, where the reply is close to known good refusals and passes. Neither trace reached a judge.
A third trace is harder. The policy tool returned eligible: partial for a damaged item, and the agent issued a partial refund with an explanation. Code cannot tell whether the amount matches the intent of the partial rule, the classifier returns 0.6, and the trace goes to the judge. That is exactly the kind of trace a judge should see.
What it costs compared with judging everything
Return to the illustrative 50,000 traces a day with three criteria. Suppose, for your agent, the tree settles decisions like this:
- Code checks settle 70% of criterion decisions.
- Semantic checks settle another 15%.
- The trained model settles another 10%.
- The judge sees the remaining 5%.
Judge cost falls from about $101,000 a month to about $5,000, plus the cost of the cheaper levels, which is small next to it. Coverage stays at 100%. Your own split will differ: an agent with lots of free-form conversation pushes more work to the upper levels, and an agent that mostly calls tools settles more in code. Measure your own split before you trust any number, including these.

Latency drops too. Most decisions land in milliseconds, so a failing trace can raise an alert while the customer is still in the conversation, not the next morning.
Keeping the tree honest
A cheap check that is wrong is worse than no check. Three habits keep the tree accurate:
- Audit settled traces. Each week, have a person review a small random set of traces each level passed or failed. Track agreement per level and per criterion.
- Watch the escalation rate. If the share of traces reaching the judge climbs, something changed: a new prompt, a new tool, new users. Look before you pay.
- Push judge findings down. When the judge keeps failing traces for the same reason, write that reason as a code or semantic check. The tree gets cheaper as it learns.
Where Trodo fits
Trodo is built around this method. It records every trace your agent produces, with spans, LLM calls, tool calls, tokens, cost and latency, and evaluates every production trace rather than a sample. Checks work as a decision tree: code checks decide first, semantic checks next, then Trodo's own models, trained for calibrated pass/fail decisions. An LLM judge only sees the traces those can't settle, which makes checking every trace affordable at a fraction of LLM-judge cost.
Failures that share a cause are grouped into issues with the failing step, the evidence and the affected users, and each issue comes with a proposed fix to the agent or to the check. Orbt checks 100% of its production traces this way and cut the time from error to shipped fix from days to under 4 hours (read the Orbt case study).
Getting started
Start with the criteria that cost you most when they fail, write the code checks for them first, and let the upper levels handle what is left. To do this on your own traces, create a free Trodo account. The Developer plan is free forever, and pricing lists what every plan includes.