Eval Analytics forAgentic Products

TwoTail makes it easy to reliably measure the quality of your agent, by keeping your evals aligned to your expertise and business outcomes.

Your agent is live. Its heart is beating.
But is it always acting intelligently?

You ran manual tests when you shipped it. Maybe you even wired up a Claude job to keep an eye on things.

But it is getting away from you. New cases keep appearing, the Claude job sends false positives, and you cannot tell what it missed.

Customers complain, and the root cause is guesswork.

And you are only checking outputs. What about outcomes? An agent that replies politely and solves nothing still passes.

TwoTail is like having an evals data analyst on your team. Always reliably measuring what is happening and why, keeping your agent quality on track.

Timothy, founder of TwoTail
Meet the founder

Your Forward Deployed Analyst

I’ve worked for some of Europe’s highest growth scaleups, and founded 3 AI startups. When you use TwoTail, you get me too. A powerful tool combined with an experienced operator figuring out the perfect way to measure and optimize your agent.

Timothy Daniell, founder Read the full story
The eval analytics loop

Four steps to measure your agent.

From the first eval you write to an always-running live suite, TwoTail has everything you need to measure your agent in production.

define annotate calibrate correlate define
  1. 01

    Define

    Bring the evals you have, build them in TwoTail from your traces, or let TwoTail propose them for you.

  2. 02

    Annotate

    Rate a smart-sampled slice of traces against your scoring rubric.

  3. 03

    Calibrate

    Judges are continuously tuned to agree with your labels, grading every trace.

  4. 04

    Correlate

    Every eval is tested against your business metrics, so you know what moves the needle.

What customers say
We used to rely on a Claude routine to find issues with our agent, but it was unreliable and I started ignoring it. Now we have evals to measure when agent quality drops - and we are confident that when it pings us on Slack it’s not a false positive.
Harold
Founder, Sliq
Sliq · outbound agent

⅓

of a founder's time reading summaries, replaced by Slack alerts he acts on

Read the case
Timothy is an expert in data modeling. With his help, we now have a clear understanding of how users behave. We highly recommend working with him.
Randy
Founder, Recharge360
Because they have so much experience in the field, they proactively find the edge cases and implement systems so those are handled correctly.
Marlon
Founder, Assembly
Eval maturity test

Where are you on the ladder?

Answer up to 6 questions to figure out where you're up to.

L1Eyeballing. You read traces by hand.
L2Labelled. You have rated a sample against a scoring rubric.
L3Calibrated. Judges continuously updated to agree with you, and you know the rate.
L4Validated. Evals are proven to predict a business metric.
L5Diagnosed. Failures traced to cause, changes measured.
Capabilities

From your labels to root cause.

Everything that makes TwoTail a complete eval analytics platform, organised to put the loop into practice.

Define

Annotate

Calibrate

Correlate

Act

Under it all

Integration

Ten minutes, and no code changes.

TwoTail reads the traces you are already producing. Point OpenTelemetry at it, or turn on the Langfuse integration and leave everything where it is.

Direct, via OpenTelemetry

Standard OTLP/JSON from any framework, straight to the TwoTail API. As easy as pasting the setup prompt into your coding agent.

Through Langfuse

Already tracing with Langfuse? Turn the integration on and TwoTail reads the same feed. No code changes, nothing moves.

Works with your stack OpenTelemetry Langfuse LangChain OpenAI Agents SDK Claude Agent SDK Vercel AI SDK
Articles

Reading on agent evals.

Guides

How to Write a Good LLM-as-a-Judge Eval

LLM-as-a-judge is a powerful technique for measuring AI quality. 10 characteristics of a good judge eval, from binary verdicts to recalibration.

Timothy Daniell · 6 min read Read the article

Common questions.

You can, and you should. Claude Code reads traces, proposes failure taxonomies, and writes eval criteria well. What it cannot do is tell you whether its judgement is right: no ground truth from you, no agreement rate, a fresh taxonomy on every rerun, and a sample rather than your whole corpus. TwoTail is what it calls to get those. Your agent does the analysis; TwoTail makes it defensible.

Online. TwoTail grades your live production traffic continuously, every trace rather than a sampled test set. Offline suites tell you whether a change broke something you already thought of. Online evals tell you what your agent is actually doing to customers today. And when you do need offline checks, you can pull your latest calibrated evals and run them there too.

Langfuse gives you somewhere to put scores. It has no opinion about which scores you should have, or whether they mean anything. That is the gap TwoTail fills. Most of our customers run TwoTail on top of Langfuse: same OTel feed, different job.

Not out of the box. In specialised fields, frontier judges agree with you as little as 64% of the time. Calibrated against a few hundred well-chosen labels, they hold up, and the agreement rate is visible before you trust a single score.

Your data is stored in isolated Postgres databases with row-level security. Each account's data is fully segregated, and all connections are encrypted in transit.

There is a free trial, and plans are priced for growth-stage teams. Full details on the pricing page.

Measure your agent properly.

Setup in 10 minutes. First calibrated judge within a week.