*This is the markdown version of this page, for agents. To see the web page, open https://trodo.ai/?view=html (or use the Human / Agent switch at the bottom)*

# Trodo: agent evaluation and observability for every trace

> Evaluating every agent trace costs a fortune. Not anymore. Trodo catches the agent failures nobody notices until your customers do, and helps you fix every one, so your agent gets better every day.

Trodo is made by Cryptique Inc. Website: https://trodo.ai. Docs: https://docs.trodo.ai. App: https://app.trodo.ai.

## What Trodo does

Trodo is a platform for teams running agents in production. It records every trace an agent produces, evaluates every one of them, groups the failures by cause and proposes the fix.

- **Trace:** every span, LLM call, tool call, token and cost, organised as a tree per trace.
- **Evaluate every trace:** checks run on 100% of production traffic, not a sample, at a fraction of the cost of sending each trace to an LLM judge.
- **Find the cause:** traces that fail the same way become one issue, with the failing step, the evidence and how many users it touched.
- **Fix:** each issue comes with a proposed change to the agent (for example its system prompt) or to the check. Approve it and it ships as a new version you can roll back in one click.
- **Lucid:** questions about your traces in plain English, answered with evidence, charts and the traces behind them.
- **Agents and alerts:** build agents that watch your data and act in Slack, Linear, GitHub and Jira; alerts when a number crosses a line.

## The problem it solves

An agent can finish a trace without an error and still do the wrong thing: refund an order outside policy, quote a number no tool returned, give a confident answer with nothing behind it. No alert fires and the trace looks normal. Teams usually find out from a customer. Checking every trace would catch these, but sending every trace to a frontier LLM judge is slow and expensive, so most teams sample a few percent and miss the rest.

## How it works

1. **Every trace gets checked.** Checks run as a decision tree. Cheap code checks decide first (did a tool call fail?), then semantic checks (is the answer grounded?), and only the traces they can't settle go to an LLM judge.
2. **Our models cost a fraction.** Hard cases go to Trodo's own models, trained for calibrated pass/fail decisions, so one request returns several checks at once at a fraction of a frontier model's cost and latency.
3. **It improves itself.** Trodo finds patterns across traces, tells your team what to fix, and sharpens the checks: a check that was too strict or too loose gets a proposed new version too.

An example: a support agent issues refunds before calling the refund-policy tool. Trodo's refund-policy check fails those traces, groups 312 of them into one issue, shows that the agent skips the policy check on damage claims, and proposes a system-prompt change. Replayed against the failing traces, the change passes; once approved it goes live as a new version.

## The loop: catch, find, recur

Most tools stop at the trace. Trodo closes the loop.

### 01 · Catch every failure
Evaluations check every trace for the failures that never throw, like a refund outside policy or an answer with nothing behind it. Alerts watch the hard failures too (tool errors, latency, cost) and fire on any evaluation, span or trace. When one trips, your team hears about it in Slack and an agent you built in Trodo gets to work.

### 02 · Find what went wrong
The agent hands the alert to Lucid. Lucid reads the failing traces, their spans and the evaluation's verdicts, and settles the question that matters: is the agent wrong, or the check? Every conclusion cites the traces behind it. If the check was wrong, Lucid proposes a new version of it for you to approve. You can also ask it anything in plain English.

### 03 · Find the recurring cause
Lucid turns failures into issues, each with its category (prompt, code, tool, infrastructure), root cause, suggested change and evidence. It merges issues that share a cause and versions its analysis as new evidence arrives. Click Fix to hand an issue to Claude Code or Cursor (connected to your repo and the Trodo MCP server), or copy it as a prompt for a local coding agent. Issues are open, merged or closed; if a closed issue's failure comes back, Trodo reopens it in triage with its full history.

Agent templates: triage alerts into issues; triage and open a GitHub issue; triage and hand to your coding agent; a weekly issue digest in Slack.

## Works with your stack

- **Agent frameworks:** LangChain, LangGraph, LlamaIndex, Vercel AI SDK, OpenAI Agents SDK, Claude Agent SDK, CrewAI, Pydantic AI, Mastra, LiveKit, OpenTelemetry, and many more.
- **Model providers:** OpenAI, Anthropic, Google Gemini, Amazon Bedrock, Azure OpenAI, Vertex AI, Mistral, Cohere, DeepSeek, Llama, Qwen, xAI, Groq, Ollama, vLLM, and many more.
- **SDKs:** Python and Node.js. One init call and one wrap around your agent; auto-instrumentation for supported frameworks. Coding agents (Claude Code, Cursor) can set it up with the Trodo skill.
- **MCP server:** query traces, issues and evaluations from Claude, Cursor or any MCP client.

## Who it is for

Teams with an agent in production, where a wrong answer costs money or trust: support and refund agents, in-product assistants, settlement and finance workflows, coding and research agents. Developers set it up in minutes; product, support and operations teams use Lucid, issues and alerts without writing queries.

## Get started

1. Sign up free at https://app.trodo.ai/signup (the Developer plan has every feature).
2. Add the Python or Node.js SDK: one init call and one wrap around your agent. Supported frameworks and model providers are instrumented automatically, or point an existing OpenTelemetry exporter at Trodo. A coding agent (Claude Code, Cursor) can do this for you with the Trodo skill or the setup prompt on the home page.
3. Traces start arriving; checks run on every one. Add your own checks or start from the defaults.
4. Turn on an alert, connect Slack, Linear, GitHub or Jira, and let an agent triage failures into issues.

Docs: https://docs.trodo.ai/start-tracing

## Everything Trodo covers

These are the same pages that exist on the site, summarised in full.

### Agent observability

Trodo records every trace your agent produces, from the first LLM call to the last tool, then checks each one and tells you which failed, why, and what to change.

Agent observability is seeing what an agent did in production: every step it planned, every model call, every tool it used, what each returned, how long it took and what it cost. It is the agent equivalent of application monitoring, built around traces and spans instead of requests.

It differs from LLM observability in scope. LLM observability looks at single model calls. An agent makes many calls, picks tools, retries, hands work to sub-agents and decides when it is done. Most real failures live in that orchestration, so agent observability follows the whole trace end to end.

Seeing a trace is only half the job. An agent can finish without an error and still give the wrong answer. Trodo checks every trace as it arrives, so the failures that never throw show up as issues instead of as customer complaints.

- **Every trace, in full:** The span tree for every trace: plans, LLM calls, tool calls, retrievals and sub-agents, with inputs, outputs, timing, tokens and cost. ([Tracing in the docs](https://docs.trodo.ai/observability/overview))
- **Checks on every trace:** Code checks, semantic checks and Trodo’s own models decide first; an LLM judge only sees what they can’t settle. Every trace scored at a fraction of the cost. ([How checks work](https://docs.trodo.ai/evaluations/overview))
- **Failures grouped, with a cause:** Traces that fail the same way become one issue, with the failing step, the evidence and how many users it touched. ([Issues in the docs](https://docs.trodo.ai/issues/overview))
- **A proposed fix:** Each issue comes with a change to the agent or to the check. Approve it and it ships as a new version you can roll back.
- **Cost and latency per trace:** Token and tool spend per trace, per agent, per user. Catch a runaway loop before the bill does. ([Cost tracking](https://docs.trodo.ai/observability/features/pricing))
- **Keep your instrumentation:** Auto-instrumentation for OpenAI, Anthropic, LangChain, LlamaIndex, the Vercel AI SDK and more, or send OpenTelemetry you already export. ([Supported frameworks](https://docs.trodo.ai/observability/features/instrumentation/frameworks/overview))

**Agent observability, LLM observability and APM**

| | Trodo | LLM tracing tools | APM |
|---|---|---|---|
| Model call traces | Yes | Yes | Partial |
| Multi-step agent traces | Yes | Partial | No |
| Tool call success, latency and cost | Yes | Yes | No |
| Every production trace checked | Yes | No | No |
| Failures grouped by cause | Yes | No | No |
| Proposed fixes, versioned | Yes | No | No |
| OpenTelemetry ingest | Yes | Partial | Yes |

#### What is agent observability?

Seeing what an agent did in production: every planned step, model call and tool call, with inputs, outputs, latency and cost, organised as traces and spans. Trodo adds a check on every trace, so failures that don’t throw an error are caught too.

#### How is agent observability different from LLM observability?

LLM observability covers single model calls. Agent observability covers the whole trace: plans, tool choices, retries and hand-offs between sub-agents, which is where most agent failures happen.

#### Does Trodo work with OpenTelemetry?

Yes. Point your existing OpenTelemetry exporter at Trodo, or use the Python or Node.js SDK, which auto-instruments the common frameworks and model providers.

#### Can Trodo run alongside Datadog or another APM?

Yes. Trodo covers the agent layer and runs alongside whatever watches your infrastructure.

#### How much does it cost?

Developer is free forever with 10,000 units and 1,000 executions a month. Pro is $249 a month with 250,000 units and 50,000 executions. Every feature is on every plan.

Full page: https://trodo.ai/agent-observability.md

### Agent analytics

Trodo finds the jobs people use your agent for, scores how well it does each one, and shows the cost, tool success and satisfaction behind every number. Ask Lucid about any of it in plain English.

Agent analytics measures how an agent performs across all of its production traffic: which tasks people give it, how often it completes them, where it fails, how long it takes and what it costs. Where observability explains one trace, analytics explains the whole agent.

The useful unit is the job, not the call. A support agent handles refunds, order status and returns, and it is rarely equally good at all three. Knowing that refunds are its most common job and its weakest one tells a team exactly where to work.

Trodo discovers those jobs from your traces, scores every trace with its checks, and lets anyone on the team ask questions about the results without writing a query.

- **Jobs, found for you:** Trodo groups your traces by what people asked the agent to do, with volume, satisfaction and failure rate for each. ([Capabilities in the docs](https://docs.trodo.ai/capabilities))
- **Every trace scored:** Checks run on every production trace, so every metric is built on all your traffic, not a sample. ([How checks work](https://docs.trodo.ai/evaluations/overview))
- **Tool success and cost:** Success rate, latency, retries and spend for every tool your agent calls. See which one drags the agent down.
- **Lucid, in plain English:** Ask Lucid why a job’s satisfaction dropped or for a weekly dashboard, and get an answer with the traces behind it. ([Lucid in the docs](https://docs.trodo.ai/lucid))
- **Reports and alerts:** Scheduled reports and alerts to Slack when a job’s failure rate or cost crosses the line you set. ([Alerts in the docs](https://docs.trodo.ai/alerts/overview))
- **Tied to users:** Every trace belongs to a user and a conversation, so you can see who a failure touched and whether they came back. ([Identify users](https://docs.trodo.ai/observability/features/users/identification))

**Agent analytics compared**

| | Trodo | LLM tracing tools | Product analytics |
|---|---|---|---|
| Jobs discovered from traces | Yes | No | No |
| Every trace scored | Yes | No | No |
| Tool success and cost | Yes | Yes | No |
| Satisfaction per job | Yes | No | No |
| Questions in plain English | Yes | Partial | Partial |
| Users and retention | Yes | No | Yes |

#### What is agent analytics?

Measuring how an agent performs across all its traffic: which jobs it is given, how often it completes them, where it fails, and what each costs. Trodo builds these numbers from checks on every trace.

#### How is agent analytics different from observability?

Observability explains a single trace. Analytics explains the agent as a whole: which jobs are growing, which are failing and which cost the most.

#### Which metrics matter most for an agent?

Task completion and satisfaction per job, failure rate, tool success rate, latency and cost per trace. Trodo tracks all of them per job and per user.

#### Do I need to write queries?

No. Ask Lucid a question in plain English and it answers with charts, tables and the traces behind them.

#### How much does it cost?

Developer is free forever with every feature: 10,000 units, 1,000 executions and $2 of Lucid credits a month. Pro is $249 a month with 250,000 units, 50,000 executions and $30 of Lucid credits.

Full page: https://trodo.ai/ai-agent-analytics.md

### Product analytics for agents

Trodo joins every checked agent trace to the user behind it, so funnels, retention and adoption show whether the agent helped, not only what the user clicked.

Classic product analytics counts clicks and page views. In a product built on an agent, the most important step often happens inside a conversation: the user asks, the agent works through tools, and the answer is right or wrong. None of that is a click.

So the questions change. Did the agent’s answer lead to the next step in the funnel? Do users whose first conversation failed come back? Which agent feature do people adopt, and which do they try once and abandon?

Trodo answers them by joining two kinds of data that usually live in separate tools: product events from your app and traces from your agent, both tied to the same user.

- **Funnels and retention:** Funnels, retention, flows and cohort comparisons over your product events. ([Product analytics in the docs](https://docs.trodo.ai/product-analytics/overview))
- **Traces joined to users:** Every trace belongs to a user and a conversation. Go from a drop in a funnel straight to the conversations behind it. ([Identify users](https://docs.trodo.ai/observability/features/users/identification))
- **Agent feature adoption:** Which jobs people bring to your agent, how often they come back, and which ones they give up on. ([Capabilities in the docs](https://docs.trodo.ai/capabilities))
- **Quality you can act on:** Every trace is checked, so a failing answer shows up next to the user it happened to, not in a support ticket a week later. ([How checks work](https://docs.trodo.ai/evaluations/overview))
- **UX health:** Rage clicks, form abandonment, errors and page performance, in the same place as the agent data.
- **Lucid, not queries:** Ask Lucid for a retention chart or a funnel by plan in plain English and get it with the data behind it. ([Lucid in the docs](https://docs.trodo.ai/lucid))

**Product analytics for agent products, compared**

| | Trodo | Classic product analytics | LLM tracing tools |
|---|---|---|---|
| Funnels and retention | Yes | Yes | No |
| Agent traces joined to users | Yes | No | Partial |
| Every trace checked | Yes | No | No |
| Agent feature adoption | Yes | Partial | No |
| Questions in plain English | Yes | Partial | Partial |

#### What is product analytics for agent products?

Measuring how people use a product built on an agent: funnels, retention and adoption, joined to what the agent did in each conversation.

#### How is it different from Mixpanel or Amplitude?

Those count events in your interface. Trodo also sees the agent’s side: each trace, whether it was right, and what it cost, tied to the same user.

#### Can I keep my existing analytics?

Yes. Trodo can run alongside your current tools; many teams start by sending agent traces and add product events later.

#### How are product events counted?

Each product event is one unit, the same as each trace and span. Developer includes 10,000 units a month and Pro 250,000, with every feature on both.

Full page: https://trodo.ai/ai-product-analytics.md

## Glossary

- **Trace:** one run of an agent from the first model call to the last tool call, stored as a tree of spans.
- **Span:** one step inside a trace: a model call, a tool call, a retrieval or a sub-agent.
- **Check (evaluation):** a test run on a trace that returns pass or fail, a score or a category. Checks are code, semantic or model-based.
- **Triage checks:** the cheap checks that run first and decide most traces, so only the unsettled ones reach a costlier model.
- **LLM judge:** a frontier model asked to grade a trace. In Trodo it only sees what cheaper checks could not settle.
- **Trodo's own models:** models trained for calibrated pass or fail decisions, which return several checks in one request at a fraction of a frontier model's cost and latency.
- **Issue:** a group of traces that fail the same way, with category, root cause, suggested change and evidence.
- **Lucid:** the part of Trodo you question in plain English. It reads traces, spans and evaluation results, answers with evidence and files issues.
- **Agent (in Trodo):** a workflow you build that watches your data and acts, for example in Slack, Linear, GitHub or Jira.
- **Alert:** a rule that fires when a number crosses a line. It can watch an evaluation, a span or a trace.
- **Unit:** one stored record (a trace, a span or a product event). Plans include a monthly number of units.
- **Execution:** one step that runs: a check step on a trace, a step in an agent you build, or an alert that fires.
- **Lucid credits:** pay for Trodo's models when Lucid and checks use them. Bring your own model keys and no credits are used.

## Documentation map

- [Traces and observability](https://docs.trodo.ai/observability/overview)
- [Start tracing](https://docs.trodo.ai/start-tracing)
- [Instrumentation guide](https://docs.trodo.ai/observability/features/instrumentation/guide)
- [Supported frameworks](https://docs.trodo.ai/observability/features/instrumentation/frameworks/overview)
- [OpenTelemetry](https://docs.trodo.ai/observability/features/instrumentation/opentelemetry)
- [Identify users](https://docs.trodo.ai/observability/features/users/identification)
- [Cost tracking](https://docs.trodo.ai/observability/features/pricing)
- [Evaluations overview](https://docs.trodo.ai/evaluations/overview)
- [Your first evaluation](https://docs.trodo.ai/evaluations/quickstart)
- [Issues](https://docs.trodo.ai/issues/overview)
- [Lucid](https://docs.trodo.ai/lucid)
- [Alerts](https://docs.trodo.ai/alerts/overview)
- [Agents](https://docs.trodo.ai/agents/overview)
- [Capabilities (jobs found in your traces)](https://docs.trodo.ai/capabilities)
- [Prompts](https://docs.trodo.ai/prompts/overview)
- [Playground](https://docs.trodo.ai/playground)
- [Reports](https://docs.trodo.ai/reports/overview)
- [MCP server](https://docs.trodo.ai/mcp)
- [Changelog](https://docs.trodo.ai/changelog)

## Pricing

Every feature is on every plan; plans differ in usage. Details: https://trodo.ai/pricing.md

- **Developer**, $0 forever: For trying Trodo on a real agent, and for side projects. Every feature: tracing, checks, issues, agents, alerts, Lucid and product analytics; 10,000 units a month; 1,000 executions a month, for checks, agents and alerts; $2 of Lucid credits a month, or bring your own model keys; 7-day retention; Up to 3 users; Community and docs support.
- **Pro**, $249 / month (or $2,499 / year): For teams with agents in production. Everything in Developer, plus: 250,000 units a month; 50,000 executions a month; $30 of Lucid credits a month, plus top-ups that never expire; 60-day retention; Unlimited users and workspaces; Email support.
- **Enterprise**, Custom: For scale, security reviews and your own deployment. Everything in Pro, plus: Units, executions, retention, rate limits and Lucid credits set to your workload; SSO with your identity provider; Self-hosted, single-tenant or BYOC deployment; Custom retention policy; DPA and security review; Dedicated Slack channel and SLA; Onboarding and team training.

A unit is one stored record (a trace, a span or a product event). An execution is one step that runs: a check step on a trace, a step in an agent you build, or an alert that fires. Lucid credits pay for Trodo's models; bring your own model keys and no credits are used. Plan credits reset each billing month; top-up credits never expire.

## Customers

Used by teams at Google, Razorpay, Airtel, Econet, Wiom, Masters Union, Zeebu and Orbt.

- **Orbt: Catching the failures that never threw.** 60% lower cost to evaluate a trace; <4 hrs error to shipped fix, down from days. "Every trace is evaluated now, at about 60% less than it cost us to judge a sample. Silent failures arrive grouped by cause, so one fix closes the whole set." (Mrudul Chauhan, Head of Product & Growth, Orbt). Story: https://trodo.ai/case-studies/orbt
- **Zeebu: Evaluating every trace, at a cost that finally works.** 80% lower cost to evaluate a trace; 100% of settlement traces evaluated, up from a sample. "Judging every trace with a frontier model cost us far too much, so we sampled. Triage checks and Trodo’s own models cut the cost of an evaluation by about 80%, and now every trace is covered." (Keshav Pandya, Chief Technical Officer, Zeebu). Story: https://trodo.ai/case-studies/zeebu
- "Trodo is not just a place to watch traces. It digs into the data with you, opens the issues, and keeps alerts, issues and checks linked, so the whole thing evolves as your product does." (Suchit Puri, Director of AI FDEs, Google)

## Security

SOC 2 Type I compliant, with Type II in progress. GDPR and CCPA compliant, DPA available. Customer data is never used to train models. Enterprise can self-host or deploy in their own cloud. Trust center: https://trodo.trust.site

## Questions

### What is Trodo?

Trodo checks every trace your agent produces in production, finds the failures nobody reports, and helps you fix them. It records each trace, scores it, groups failures into issues with their cause, and proposes the change that fixes them.

### How is Trodo different from a tracing or observability tool?

A tracing tool shows you what your agent did. Trodo also tells you whether it was right. Every trace is checked, failures are grouped and explained, and each one comes with a proposed fix to the agent or to the check.

### Can I evaluate every trace without the cost getting out of hand?

Yes. Checks run as a decision tree: cheap code checks decide first, then semantic checks, then our own models trained for these decisions. A frontier LLM judge only sees the traces the others can’t settle, so most traces are scored at a fraction of the cost.

### Can’t I paste my traces into Claude or ChatGPT instead?

For a single trace, yes. In production there are thousands, most of them repetitive. Trodo does the narrowing: it finds the failing traces, the step where each one went wrong and the evidence, and you can ask about them from Claude or Cursor through the Trodo MCP server.

### How long does setup take?

Minutes. Install the Python or Node.js SDK, add one init call and one wrap around your agent. Or ask your coding agent to do it with the Trodo skill. If you already export OpenTelemetry, point it at Trodo.

### Does Trodo change my agent on its own?

No. Every fix is a proposal. You review the cause and the change, approve it, and it goes live as a new version you can roll back in one click.

## Talk to us

- Get started free: https://app.trodo.ai/signup
- Book 20 minutes with the founders: https://trodo.ai/#book
- Email: hello@trodo.ai
- X: https://x.com/Trodo_ai · LinkedIn: https://www.linkedin.com/company/trodoai

## More pages

- Pricing: https://trodo.ai/pricing.md
- Customers: https://trodo.ai/customers.md
- Agent observability: https://trodo.ai/agent-observability.md
- Agent analytics: https://trodo.ai/ai-agent-analytics.md
- Product analytics for agents: https://trodo.ai/ai-product-analytics.md
- Blog: https://trodo.ai/blog.md
- Docs: https://docs.trodo.ai