Skip to content

Conversation evals

Adversarial evals prove your agent can’t be broken. Conversation evals answer the everyday question: are its answers actually good?

  1. You provide an agent endpoint and a set of test cases — each is a prompt (or multi-turn conversation) for the agent.
  2. The eval sends every test case to your agent and captures its responses.
  3. A judge model scores each response against your chosen criteria.
  4. You get per-test scores, judge reasoning, and conversation transcripts, plus an overall pass rate against your threshold.

The default criteria set is accuracy, helpfulness, safety, relevance, conciseness — the five axes a human reviewer would instinctively apply. You can narrow the list, and you can add custom criteria with your own definitions when your product has domain-specific quality bars (“cites the pricing page correctly”, “never recommends a competitor”).

The pass threshold (0–1, default 0.7) decides what counts as a passing test. Raise it for high-stakes surfaces; lower it while iterating on prompts.

  • Per-test scores on every criterion, with the judge’s written reasoning — not just numbers, so you can disagree intelligently.
  • Full transcripts of what the agent actually said, attached to each verdict.
  • Overall pass rate for the run, comparable across runs — the number to graph over time.
  • Treat the suite as regression coverage for prompts: a model or system-prompt change that drops pass rate is a caught bug, not a mystery.
  • Skim the judge’s reasoning on failures — some will be judge artifacts; the fix there is a sharper criterion definition, not a blind prompt rewrite.
  • Keep a small set of golden conversations you hand-checked; they anchor the judge’s calibration.

Adversarial safety is covered separately by the adversarial suite; the two compose — safety is also a default judge criterion here.