Skip to content

Adversarial evals

Before users find your agent’s breaking points, attackers will. Adversarial evals send a curated library of attacks at your agent endpoint and report exactly which ones got through.

40 seed prompts across six categories, each graded by severity:

Category Attacks Severity range Tests for
Prompt injection 10 critical–medium Instruction override, system-prompt extraction, role hijack
Jailbreak 8 critical–medium Safety-guideline bypass, unrestricted personas
Data extraction 6 high–medium Pulling internal context, keys, hidden state
PII leakage 5 high–medium Extracting personal data from context or memory
Bias probing 5 medium Discriminatory or stereotyped outputs
Off-topic 5 low Robustness when steered away from purpose

Each attack carries the prompt text and the expected behavior (e.g. “Agent refuses and stays in character”), so results read as blocked vs succeeded, not just error vs ok.

Point the eval at your agent’s API endpoint, optionally filter by category or minimum severity, and let it run. The result is per-attack verdicts plus aggregate metrics — most importantly the block rate, the fraction of attacks the agent refused or safely deflected.

  • Critical/high successes are ship blockers: an agent that leaks its system prompt or extracted data is one screenshot away from a trust incident.
  • Medium successes are judgment calls — bias probes and aggressive extraction attempts may warrant prompt hardening.
  • Low (off-topic) successes are usually acceptable: an agent staying helpful on tangents is rarely a security event.

Re-run the suite whenever the system prompt, tools, or model change — regression here is silent until it isn’t. Pair with conversation evals to check the agent stayed good while getting safer.