Conversation evals
Adversarial evals prove your agent can’t be broken. Conversation evals answer the everyday question: are its answers actually good?
How an eval run works
Section titled “How an eval run works”- You provide an agent endpoint and a set of test cases — each is a prompt (or multi-turn conversation) for the agent.
- The eval sends every test case to your agent and captures its responses.
- A judge model scores each response against your chosen criteria.
- You get per-test scores, judge reasoning, and conversation transcripts, plus an overall pass rate against your threshold.
Judge criteria
Section titled “Judge criteria”The default criteria set is accuracy, helpfulness, safety, relevance, conciseness — the five axes a human reviewer would instinctively apply. You can narrow the list, and you can add custom criteria with your own definitions when your product has domain-specific quality bars (“cites the pricing page correctly”, “never recommends a competitor”).
The pass threshold (0–1, default 0.7) decides what counts as a passing test. Raise it for high-stakes surfaces; lower it while iterating on prompts.
What you get back
Section titled “What you get back”- Per-test scores on every criterion, with the judge’s written reasoning — not just numbers, so you can disagree intelligently.
- Full transcripts of what the agent actually said, attached to each verdict.
- Overall pass rate for the run, comparable across runs — the number to graph over time.
Using evals well
Section titled “Using evals well”- Treat the suite as regression coverage for prompts: a model or system-prompt change that drops pass rate is a caught bug, not a mystery.
- Skim the judge’s reasoning on failures — some will be judge artifacts; the fix there is a sharper criterion definition, not a blind prompt rewrite.
- Keep a small set of golden conversations you hand-checked; they anchor the judge’s calibration.
Adversarial safety is covered separately by the adversarial suite; the two compose — safety is also a default judge criterion here.