Eval Analytics for Agent Builders

TwoTail makes it easy to reliably measure your agent, by keeping it aligned to your expertise and business outcomes.

Works with your stack OpenTelemetry Langfuse LangChain OpenAI Agents SDK Claude Agent SDK Vercel AI SDK

Your agent is live. Its heart is beating.
But is it always acting intelligently?

You ran manual tests when you shipped it. Maybe you even wired up a Claude job to keep an eye on things.

But it is getting away from you. New cases keep appearing, the Claude job sends false positives, and you cannot tell what it missed.

Customers complain, and the root cause is guesswork.

And you are only checking outputs. What about outcomes? An agent that replies politely and solves nothing still passes.

TwoTail is like having an evals data analyst on your team. Always reliably measuring what is happening and why, keeping your agent quality on track.

The eval analytics loop

Four steps to measure your agent.

From the first eval you write to an always-running live suite, TwoTail has everything you need to measure your agent in production.

define annotate calibrate correlate define
  1. 01

    Define

    Bring the evals you have, build them in TwoTail from your traces, or let TwoTail propose them for you.

  2. 02

    Annotate

    Rate a smart-sampled slice of traces against your scoring rubric.

  3. 03

    Calibrate

    Judges are continuously tuned to agree with your labels, grading every trace.

  4. 04

    Correlate

    Every eval is tested against your business metrics, so you know what moves the needle.

Eval maturity test

Where are you on the ladder?

Answer up to 6 questions to figure out where you're up to.

L1Eyeballing. You read traces by hand.
L2Labelled. You have rated a sample against a scoring rubric.
L3Calibrated. Judges continuously updated to agree with you, and you know the rate.
L4Validated. Evals are proven to predict a business metric.
L5Diagnosed. Failures traced to cause, changes measured.
Capabilities

From your labels to root cause.

Everything that makes TwoTail a complete eval analytics platform, organised to put the loop into practice.

Define

Annotate

Calibrate

Correlate

Act

Under it all

Integration

Ten minutes, and no code changes.

TwoTail reads the traces you are already producing. Point OpenTelemetry at it, or turn on the Langfuse integration and leave everything where it is.

Direct, via OpenTelemetry

Standard OTLP/JSON from any framework, straight to the TwoTail API. As easy as pasting the setup prompt into your coding agent.

Through Langfuse

Already tracing with Langfuse? Turn the integration on and TwoTail reads the same feed. No code changes, nothing moves.

Common questions.

You can, and you should. Claude Code reads traces, proposes failure taxonomies, and writes eval criteria well. What it cannot do is tell you whether its judgement is right: no ground truth from you, no agreement rate, a fresh taxonomy on every rerun, and a sample rather than your whole corpus. TwoTail is what it calls to get those. Your agent does the analysis; TwoTail makes it defensible.

Online. TwoTail grades your live production traffic continuously, every trace rather than a sampled test set. Offline suites tell you whether a change broke something you already thought of. Online evals tell you what your agent is actually doing to customers today. And when you do need offline checks, you can pull your latest calibrated evals and run them there too.

Langfuse gives you somewhere to put scores. It has no opinion about which scores you should have, or whether they mean anything. That is the gap TwoTail fills. Most of our customers run TwoTail on top of Langfuse: same OTel feed, different job.

Not out of the box. In specialised fields, frontier judges agree with you as little as 64% of the time. Calibrated against a few hundred well-chosen labels, they hold up, and the agreement rate is visible before you trust a single score.

Your data is stored in isolated Postgres databases with row-level security. Each account's data is fully segregated, and all connections are encrypted in transit.

There is a free trial, and plans are priced for growth-stage teams. Full details on the pricing page.

Measure your agent properly.

Setup in 10 minutes. First calibrated judge within a week.

Timothy, founder of TwoTail

Want it done with you? I run the first calibration alongside your team: error analysis, taxonomy, judges. Timothy, founder