AI Validation
AI output needs the same testing discipline as shipped code — more, because it fails quietly. This section covers the four QualityMax validation surfaces: attacking your agent on purpose, scoring its conversations, separating generated-test hallucinations from real bugs, and sending AI personas through your product like real users.
The four surfaces
Section titled “The four surfaces”| Surface | Question it answers | Tool |
|---|---|---|
| Adversarial evals | Can your agent be broken, hijacked, or leaked? | Curated attack suite, 40 seed prompts across 6 categories |
| Conversation evals | Are your agent’s answers actually good? | Multi-criteria judge scoring with pass thresholds |
| Hallucination gate | Did the generated test test something real? | Execution-based verification states |
| Agentic Eyes | Can a real user succeed on your site? | Persona-driven browser reviews and model-consensus panels |
Why execution beats inspection
Section titled “Why execution beats inspection”Every surface here shares one design rule: verify against reality, not the model’s word for it. The hallucination gate classifies a generated test by running it against the live target and reading the real status codes — a test that asserts against an endpoint that doesn’t exist gets quarantined no matter how plausible the code looks. Agentic Eyes personas drive a real browser and report what actually happened on the journey, not what a static review guesses would happen.
Where each fits
Section titled “Where each fits”- Building an AI agent or chatbot → start with adversarial evals, then conversation evals for quality regression.
- Generating tests with AI → the hallucination gate runs automatically in the pipeline; read it to understand quarantine and verification badges.
- Improving conversion or onboarding → Agentic Eyes gives you a non-QA buyer’s report on your own site.