Trace Sampling vs Full Coverage in Agent Observability
A 5% sample misses a failure that happens 10 times a day on about 60% of days. Here is the math, when sampling is fine, and how triage makes checking every trace affordable.
Many teams evaluate a small slice of production: 1 to 5% of traces, scored by an LLM judge or a human reviewer. The dashboards look fine. Then a customer screenshot shows the agent doing something expensive that no sampled trace ever showed.
That isn't bad luck. It is what sampling does. A sample is good at estimating how often common things happen and bad at finding rare things at all. In agents, the rare failures are often the costly ones: a refund that breaks policy, a wrong account closed, a confident answer with no source behind it.
This article works through the probability of missing a failure under trace sampling, shows why agent failures follow a long tail, explains when sampling is the right call, and shows how triage makes checking 100% of traces affordable.
What trace sampling means in agent observability
In agent observability, sampling can happen at two points. You can sample what you record, keeping only some traces at all. Or you can record everything and sample what you evaluate, scoring only a fraction with checks. The second is more common for agents, because recording is cheap and evaluation with an LLM judge is not.
Both kinds share the same blind spot. A trace that is never checked cannot raise a failure. The question is how many failures fall into that gap.
The math: how many failures a 5% sample sees
Say your agent handles 50,000 traces a day and you check a random 5% of them. That is 2,500 checked traces a day. These numbers are illustrative; plug in your own.
A failure affecting 0.2% of traces
A failure that hits 0.2% of traces happens 50,000 × 0.002 = 100 times a day. In a 5% sample you expect to see 2,500 × 0.002 = 5 of them. The other 95 go unchecked. At a 1% sample you see about 1 a day and miss 99.
Five a day is enough to notice, so sampling does find this one. But you see 5% of the evidence, and the 95 affected users never show up in your failure review.
A rarer failure: 10 times a day
Now take the refund agent failure: it refunds an order outside the 30-day policy before calling the policy tool. Say it happens 10 times a day, or 0.02% of traces. Each of those traces has a 5% chance of being sampled, so the chance that none of today's 10 is checked is:
0.95^10 ≈ 0.60
So on 60% of days, this failure is invisible. Over a full week, 70 occurrences, the chance of missing all of them drops to 0.95^70 ≈ 0.03. That sounds reassuring until you look at what it means in practice:
- You expect to see 0.5 of these failures a day, or one every 2 days.
- One stray failing trace rarely looks like a pattern. To collect 5 examples, enough to see that they share a cause, takes about 10 days at 5% and about 50 days at 1%.
- At a 1% sample, the chance of seeing none of the day's 10 is 0.99^10 ≈ 0.90. Even across a whole week, the chance of missing all 70 is 0.99^70 ≈ 0.49, a coin flip.
If each wrong refund costs $80, those 10 a day cost $800 a day, about $5,600 a week, while the sample shows almost nothing.

The long tail of agent failures
Agent failures are not one big problem. They are a few common modes and a long list of rare ones. An illustrative day for a support agent with 50,000 traces might look like this:
- One common mode, such as a tool timeout: 250 traces a day. A 5% sample sees about 12.
- One moderate mode, such as answers that ignore retrieved documents: 100 a day. The sample sees about 5.
- A handful of modes at 5 to 20 a day, such as the refund-before-policy failure. For each one, the sample sees between one a day and one every four days.
- Dozens of modes at under 1 a day: a wrong currency, a closed account, a promise the company can't keep. The sample almost never sees them.
Sampling concentrates your attention on the head of that list, which is usually the part you already know about. The tail is where new and costly problems live, and it is also where each individual mode is too rare to show up in a sample. Added together, the tail can affect more users than the head. Many of these are silent failures: the trace returns status OK, looks fast and normal, and only a check on the content shows it was wrong.
When sampling is fine
Sampling is a good tool for the right question. It works well when:
- You are estimating a common rate. If 5% of traces fail a check, a random sample of 2,500 estimates that rate to within about ±0.85 percentage points at 95% confidence. You don't need all 50,000.
- You are tracking trends and aggregates: latency percentiles, token cost per trace, overall pass rate week over week.
- The check is very expensive, such as human review. You cannot have a person read 50,000 traces a day, so a well-chosen sample is the only option.
- Failures are cheap and frequent, so missing some has little cost and the ones you do see are representative.
Sampling breaks down when the goal is to find failures rather than measure them, and when a single miss is expensive. That describes most policy, safety and money-moving behavior in agents.
Make any sample smarter
If you must sample, don't sample uniformly. Oversample traces with a signal: a tool error, a refund, a negative rating, an unusually long trace, a new user. But notice that deciding which traces carry a signal already requires a cheap check on every trace. That is triage, and once you have it, full coverage is close.
How triage makes 100% coverage affordable
Sampling exists because judging every trace with a frontier LLM is expensive. The fix is not a smaller sample. It is to stop sending every trace to the most expensive check. Most traces can be settled by something much cheaper.
- Code checks first. Did a tool call fail? Was the refund tool called before the policy tool? Did the reply contain an order number that exists? These are free and deterministic.
- Semantic checks next. Is the answer grounded in the retrieved documents? Does it contradict the policy text? Embedding and small-model checks are cheap per trace.
- An LLM judge last, only for the traces the earlier steps can't decide.
Here is an illustrative cost comparison. At 50,000 traces a day, you have about 1.5 million traces a month. If an LLM judge costs $0.01 per trace, judging every trace costs about $15,000 a month, and a 5% sample costs $750. Now suppose that in your own setup, code checks settle 80% of traces and semantic checks settle another 15%. The judge sees the remaining 5%, about 75,000 traces, for roughly $750 a month, plus the small cost of the cheaper tiers. You pay about what the sample cost and check every trace. For a closer look at judge cost and accuracy, see LLM-as-a-judge cost and accuracy.
Where Trodo fits
Trodo is built on this idea. It records every trace, with spans, LLM calls, tool calls, tokens, cost and latency, and evaluates every production trace, not a sample. Checks work as a decision tree: cheap code checks decide first, then semantic checks, then Trodo's own models trained for calibrated pass/fail decisions. An LLM judge only sees the traces those can't settle, which keeps full coverage at a fraction of LLM-judge cost.
Because every trace is checked, the rare refund failure shows up on the first day it happens, not on day ten. Failures that share a cause are grouped into issues with the failing step, the evidence and the affected users, and each issue comes with a proposed fix to the agent's prompt or to the check. Orbt checks 100% of its production traces and cut the time from error to shipped fix from days to under 4 hours.
Getting started
Take your current sample rate and your daily trace volume, and compute (1 − s)^n for a failure you care about. If the answer is uncomfortable, list the rules you can express as code checks and apply them to every trace first. Keep the judge for what is left. If you already have a regression dataset, the failures you find this way are the best new cases for it; see offline vs online evaluation.
To check every trace without building the triage yourself, sign up for Trodo. The Developer plan is free forever with every feature included; pricing lists the limits for each plan.