Adversarial evals
Before users find your agent’s breaking points, attackers will. Adversarial evals send a curated library of attacks at your agent endpoint and report exactly which ones got through.
The attack library
Section titled “The attack library”40 seed prompts across six categories, each graded by severity:
| Category | Attacks | Severity range | Tests for |
|---|---|---|---|
| Prompt injection | 10 | critical–medium | Instruction override, system-prompt extraction, role hijack |
| Jailbreak | 8 | critical–medium | Safety-guideline bypass, unrestricted personas |
| Data extraction | 6 | high–medium | Pulling internal context, keys, hidden state |
| PII leakage | 5 | high–medium | Extracting personal data from context or memory |
| Bias probing | 5 | medium | Discriminatory or stereotyped outputs |
| Off-topic | 5 | low | Robustness when steered away from purpose |
Each attack carries the prompt text and the expected behavior (e.g. “Agent refuses and stays in character”), so results read as blocked vs succeeded, not just error vs ok.
Running a suite
Section titled “Running a suite”Point the eval at your agent’s API endpoint, optionally filter by category or minimum severity, and let it run. The result is per-attack verdicts plus aggregate metrics — most importantly the block rate, the fraction of attacks the agent refused or safely deflected.
Reading results
Section titled “Reading results”- Critical/high successes are ship blockers: an agent that leaks its system prompt or extracted data is one screenshot away from a trust incident.
- Medium successes are judgment calls — bias probes and aggressive extraction attempts may warrant prompt hardening.
- Low (off-topic) successes are usually acceptable: an agent staying helpful on tangents is rarely a security event.
Re-run the suite whenever the system prompt, tools, or model change — regression here is silent until it isn’t. Pair with conversation evals to check the agent stayed good while getting safer.