LLM Hallucination Detection for Agents: Grounding Checks
How to detect hallucinations in agents: check that every number and claim traces back to a tool output in the same trace, then layer semantic and model-based checks on top.
A chat model that hallucinates gives a wrong answer. An agent that hallucinates gives a wrong answer that looks verified. It tells the customer "your order shipped on March 3" when the order tool returned no ship date. It says "refunds are allowed within 60 days" when the policy tool said 30.
Agents also make hallucinations easier to catch than plain chat does. Every fact the agent should rely on passed through the trace as a tool output or a retrieved document. So the core question becomes close to mechanical: does each claim in the answer trace back to something in the same trace?
This article covers the checks that catch most agent hallucinations: grounding, citation, consistency and abstention. It shows how to stack cheap string and number matching under semantic and model-based checks so you can check every trace, and how to handle false positives so your team keeps trusting the results.
Why agent hallucinations are different
LLM evaluation for factual accuracy usually needs an outside reference: a labeled answer or a search result. For an agent, most of the reference is already inside the trace. That changes what a hallucination looks like. There are three kinds worth checking for:
- Fabricated facts: a number, date, name or ID that appears in no tool output. "Refund of $84.20 issued" when the order total was $48.20.
- Misread facts: the value exists in the trace but is attached to the wrong thing. The agent quotes the 60-day exchange window from the shipping document as the refund window.
- Fabricated actions: the agent says it did something it did not do. "I've issued your refund" with no issue_refund span in the trace.
Fabricated actions are the most expensive, because customers act on them. They are also a classic silent failure: the trace finishes fast with a status of OK, and nothing looks wrong until the customer writes back.
Grounding checks: does every claim trace back to the trace?
A grounding check extracts claims from the final answer and looks for support in the tool outputs and retrieved documents from the same trace. Start with the claims that are cheapest to verify and most costly to get wrong: numbers, dates, IDs, amounts and names.
Layer 1: exact matching for numbers and identifiers
Pull every number, currency amount, date and ID from the answer with regular expressions. Normalize them, so that "$48.20", "48.2" and "USD 48.20" compare equal, as do "3 March" and "2026-03-03". Then look for each normalized value in the evidence.
The derivable step matters. Agents compute. "You have 12 days left to return this" comes from the delivery date and the policy length, not from any tool output verbatim. Allow simple arithmetic over grounded values before you fail a trace.
Layer 2: semantic matching for statements
Not every claim is a number. "Your plan includes priority support" needs a comparison against the plan document. Split the answer into short, single-fact claims. For each claim, find the closest evidence passages in the trace by embedding similarity, then score entailment: does the passage support the claim, contradict it, or say nothing? A contradicted claim is a strong failure. A claim with no supporting passage above your threshold is a candidate for the next layer.
Layer 3: a model decides what the first two can't
Some claims need careful reading, such as a paraphrased policy with a condition attached. Give a model the claim, the top three evidence passages and a narrow question: "Is this claim supported by these passages? Answer supported, contradicted or not found." Keep the question narrow. A judge asked "is this answer accurate?" drifts. A judge asked about one claim against three passages is far more consistent.
Citation and consistency checks
Citation checks
If your agent cites sources, such as help center links, ticket IDs or knowledge base article numbers, check citations separately from grounding:
- Existence: every cited ID or URL appears in a retrieval or tool span in this trace. A citation to a document the agent never retrieved is fabricated. This is a pure code check.
- Support: the cited passage actually supports the sentence it is attached to. This is a semantic or model-based check, scoped to one sentence and one passage.
- Coverage: factual sentences carry a citation, if your product requires one.
Existence failures are cheap to detect and unambiguous, which makes them a good first alert.
Consistency checks
Consistency catches hallucinations that grounding misses because the evidence is ambiguous or absent:
- Answer vs actions: if the answer says "I've refunded you", there must be a successful issue_refund span. This is a code check on the span tree, and it catches fabricated actions directly.
- Answer vs itself: two different amounts for the same refund in one reply.
- Across turns: the agent says 30 days in turn two and 60 days in turn five of the same conversation.
- Across samples: generate the answer several times for the same input. Claims that change between samples are more likely fabricated. This is useful when replaying traces offline and usually too costly for every production trace.
Abstention checks: does the agent admit it doesn't know?
The right response to missing evidence is "I couldn't confirm that" or a handoff to a person. An abstention check finds traces where the evidence was thin and asks whether the agent said so. It fails a trace when two conditions are both true:
- Missing evidence: a tool returned empty or errored, or retrieval scores were low. A code check on spans.
- A confident answer anyway: the reply states a specific fact with no hedge and no escalation. A semantic or model-based check.
Take the refund agent. get_refund_policy times out, and the agent tells the customer "you're within the 30-day window, your refund is on its way." Nothing in the trace supports it. The right behavior is to say the policy could not be checked and escalate.
Pair this with the inverse check. An agent that refuses when the evidence was clearly present is failing in a different way, and prompts tuned hard against hallucination tend to produce exactly that. Track both, or you will trade one problem for the other.
Stacking the checks: cheap first, judges last
Sending every claim of every trace to an LLM judge is slow and expensive. Order the checks so each layer only sees what the layer before it couldn't decide:

- Code checks: citation existence, answer-vs-action consistency, number and ID matching, tool errors followed by confident answers. Fast, deterministic and nearly free.
- Semantic checks: embedding similarity and entailment for statements. Cheap, and they handle paraphrase.
- Trained models: calibrated pass/fail decisions on the claims still open.
- An LLM judge: only for the traces still undecided, with a narrow question per claim.
A worked example. Say the refund agent handles 30,000 traces a day. Most answers contain an order ID, an amount and a date, and each of those either matches a tool output or it doesn't. Suppose code checks settle 70% of traces, semantic checks settle most of the rest, and a few percent reach a judge. You are judging hundreds of traces a day instead of 30,000, and they are the hard ones. The percentages are illustrative, but the shape holds for most tool-heavy agents. LLM-as-a-judge cost and accuracy works through the cost math.
Handling false positives
A hallucination check that cries wolf gets muted. These are the usual sources, with the fix for each:
- Derived values: dates and totals computed from evidence. Allow arithmetic over grounded values.
- Formatting: "$1,200" vs "1200.00", "Mar 3" vs "2026-03-03". Normalize before matching.
- General knowledge: "business days exclude weekends" needs no tool output. Only check claims about the customer, the order and the policy, or keep an allowlist.
- Evidence outside tool spans: facts in the system prompt and earlier turns are valid sources. Include them as evidence.
- Borderline judge decisions: a claim the judge flips on between calls. Route it to uncertain rather than fail, and have a person label a sample.
Measure precision per check: of the traces a check failed, how many would a person agree were hallucinations? Label 50 failed traces for each new check before you alert on it. If precision is low, fix the check before you touch the agent. Every labeled disagreement is also data for the next version of the check.
Where Trodo fits
Trodo records every trace with its spans, tool calls and LLM calls (agent observability), so the tool outputs and documents you ground against are already captured. It evaluates every production trace as a decision tree: cheap code checks decide first, then semantic checks such as "is the answer grounded?", then Trodo's own models trained for calibrated pass/fail decisions. An LLM judge only sees the traces those can't settle, which keeps checking every trace at a fraction of LLM-judge cost.
Hallucinations that share a cause are grouped into issues with the failing step, evidence and affected users. Each issue comes with a proposed fix to the agent or to the check, which is how a noisy grounding check gets corrected instead of ignored. You can also put plain-English questions to your traces with Ask, and get answers with evidence and the traces behind them.
Getting started
Pick the claim type that would hurt most if it were wrong, usually amounts and dates, and add a number grounding check plus an answer-vs-action check. Check every trace, label the first 50 failures, and tune before adding semantic and model-based layers. Sign up for the free Developer plan, or compare plans on pricing.