LLM as a Judge: Accuracy, Bias and Cost in LLM Evaluation
An LLM judge can grade agent output at scale, but it has biases, a real cost and needs calibration. Here is how it works, where it holds up and when cheaper checks do the job better.
Using one LLM to grade another is now the default way teams score open-ended output. It is quick to set up: write a rubric, send the output, get a verdict. It also fails in ways that are easy to miss, and at production volume it is often the largest line in the evaluation budget.
This article explains how LLM-as-a-judge works, where it is reliable, the biases that skew it, how to calibrate a judge against human labels, and what it costs at scale. It ends with a simple rule for when a cheaper check should do the work instead.
How LLM-as-a-judge works
A judge is a model call with three parts: the thing being graded, a rubric that says what good looks like, and an output format the rest of your pipeline can parse. There are three common setups.
- Direct grading: the judge reads one output and returns pass/fail or a score against the rubric. This is the most common setup for production traces.
- Pairwise comparison: the judge sees two outputs for the same input and picks the better one. Useful when comparing prompt versions.
- Reference-based grading: the judge compares the output to a known good answer or to source material, such as tool outputs or retrieved documents.
A minimal direct-grading judge for a customer-support agent that issues refunds looks like this:
Two details matter more than the model choice. Ask for a binary verdict with a short reason, not a 1 to 10 score. And give the judge the evidence, here the tool outputs, not just the final reply.
Where an LLM judge is reliable
Judges do well when the question is narrow and the evidence is in front of them:
- Binary questions with a clear rubric: did the reply follow this written policy?
- Grounding against supplied context: is each claim supported by these documents?
- Tone and format rules that a person could check in seconds.
- Comparing two outputs where one is clearly better.
They do poorly on:
- Fine-grained scales. The gap between a 6 and a 7 is mostly noise, and the same trace can get both on two calls.
- Facts the judge has to know rather than read, such as current prices or internal rules left out of the prompt.
- Correctness that needs execution, like whether generated SQL returns the right rows.
- Long traces where the key evidence is buried among dozens of spans.
- Counting and date math, like whether day 31 falls outside a 30-day window. One line of code does this better.
Known biases in LLM judges
Judges are not neutral readers. Research on LLM-as-a-judge has documented several consistent biases. Knowing them lets you design around them.
Position bias
In pairwise comparison, judges tend to favor one position, often the first answer shown. The fix is cheap: grade each pair twice with the order swapped, and count a win only when both orders agree. Treat disagreement as a tie.
Verbosity bias
Judges tend to rate longer, more detailed answers higher, even when the extra length adds nothing or hides an error. For a support agent this rewards long, apologetic replies over short, correct ones. State in the rubric that length is not a criterion, and test the judge on pairs where the shorter answer is the correct one.
Self-preference
A judge can rate outputs from its own model family higher, likely because they match its own style. If your agent and your judge share a model, use a judge from a different family, or at least check agreement with human labels on outputs from both.
Trusting the agent's own claims
This one matters most for agents. If the reply says I checked our policy and you qualify, a judge reading only the reply tends to believe it. In the refund case, the judge passes a refund for an order delivered 41 days ago because the agent said it checked. Give the judge the tool calls and their outputs, and tell it to rely on evidence, not on what the agent says about itself. This is a common source of silent agent failures: the trace looks fine and the verdict agrees.
Calibrating a judge against human labels
A judge you have not measured is an opinion. Calibration tells you how often it agrees with people who know the domain, and where it goes wrong. Here is a method that fits in a week:
- Pull 200 to 300 production traces for one criterion. Over-sample the kinds you care about: refunds, escalations, long conversations.
- Have two people label each trace pass or fail, independently, with the same rubric the judge will get.
- Measure agreement between the two people first. If they agree on only 80% of traces, the rubric is unclear. Fix the rubric before you blame the judge.
- Settle disagreements, then grade the same traces with the judge.
- Compute precision and recall on the fail class. Precision tells you how many flagged traces are real failures. Recall tells you how many real failures the judge caught.
- Read every disagreement. Most point to a rubric gap, missing evidence or one of the biases above.
- Keep a held-out set of labeled traces and re-check the judge whenever you change its prompt, its model or the agent.
Worked example
Say you label 250 refund traces and your reviewers agree that 40 are failures. The judge flags 52 traces. Of those, 34 are among the 40 real failures.
- Precision: 34 / 52 = 65%. About one in three flags is noise.
- Recall: 34 / 40 = 85%. Six real failures slip through.
Reading the 18 false flags shows most are long, apologetic refusals the judge read as policy breaks. Reading the 6 misses shows the agent claimed it had checked the policy. Adding the tool outputs to the prompt and one rubric line about evidence fixes most of both. Then you measure again, on the held-out set.
Cost and latency at scale
A judge call on an agent trace is not small. The judge has to read the conversation, the tool calls and the rubric, then write its reasoning.
An illustrative calculation: say your agent handles 20,000 traces a day and a judge call reads 5,000 input tokens and writes 250. At an example price of $3 per million input tokens and $15 per million output tokens, one call costs about $0.019. That is about $375 a day, or roughly $11,000 a month, for a single criterion. Four criteria, each judged separately, come to about $45,000 a month. Pairwise grading with swapped order doubles the calls.

Latency is the second cost. Each judge call takes seconds, so a judge can't sit in front of a live response without slowing the agent down. Judging asynchronously works, but at peak traffic the queue grows, and a failure is found minutes or hours after the customer saw it.
The usual reaction is to judge a sample. That controls the bill and gives up on rare failures, which are often the expensive ones.
When to use cheaper checks instead
The rule is simple: use the cheapest check that can decide the question reliably, and keep the judge for what is left.

- Code for anything structural: tool errors, schema, whether check_refund_policy was called before issue_refund, whether the order date falls inside 30 days.
- Semantic checks for meaning that stays close to the text: similarity to known failures, grounding overlap with tool outputs, entailment with a small NLI model.
- A small trained model for criteria you check on every trace and have labels for. Trained on a few hundred human-labeled traces, it gives calibrated pass/fail probabilities at a fraction of a judge's cost.
- The LLM judge for ambiguous cases that need reasoning over context: a partial refund for a damaged item, or a reply that is technically correct but misleading.
Arranged as a decision tree, each level settles what it can and passes the rest up. Most traces never reach the judge, so you can afford to check all of them. We walk through building one in how to evaluate every agent trace.
Where Trodo fits
Trodo applies this approach to every production trace. It records each trace with its spans, LLM calls, tool calls, tokens, cost and latency, then evaluates it with a decision tree: code checks first, then semantic checks, then Trodo's own models, trained for calibrated pass/fail decisions. An LLM judge only sees the traces those can't settle, so checking every trace costs a fraction of judging them all.
Failures that share a cause become issues with the failing step, evidence and affected users. Each issue comes with a proposed fix: a change to the agent, such as its system prompt, or to the check itself when the check was wrong. For the wider picture of what to watch in production, see agent observability.
Getting started
Pick one criterion, label 200 traces, and measure your judge before you trust it. If you want every trace checked without a judge bill to match, sign up for Trodo. The Developer plan is free forever, and pricing shows what each plan includes.