Eval Builder
Start from the evals you already have, or write new ones by labelling the traces they will grade. Every eval is built on a scoring rubric: what good looks like, written down. Severity runs on anchored scales with behavioural anchors, so two people reading the rubric score the same trace the same way, and the headline number is binary at the threshold you choose.
evals · support agent