Hallucination gate
An AI-generated test can look perfect and test nothing. The hallucination gate is the pipeline stage that tells the difference — by executing the generated test against the live target and classifying the result.
Why it exists
Section titled “Why it exists”The founding incident: a dogfood crawl against a real product generated 40 test scripts, and 38 of them asserted against an API contract that didn’t exist — every create call returned 400 unknown action. Every test “passed review” as code. All of them were hallucinations. Static inspection cannot catch this class of failure; only running against reality can.
The four verdicts
Section titled “The four verdicts”| State | Meaning | What happens |
|---|---|---|
| Verified | Ran green against the live target | Counts as automation |
| Failed live | Target exists, behaves wrong | A real found bug — keep it. A failing test that reproduces a genuine defect is a legitimate test |
| Hallucinated | Asserted endpoint/contract/selector doesn’t exist | Quarantined — never counts as the test case’s automation |
| Inconclusive | Infra/transient failure (timeout, 502–504, connection refused) | Badge only, no verdict |
(A fifth state, unverified, marks scripts never executed — e.g. syntax-only saves.)
The distinction that matters: a hallucination is the only state that must never count as coverage. failed_live is gold — an automated reproduction of a real bug. hallucinated is debt wearing a test’s clothes.
Deterministic-first design
Section titled “Deterministic-first design”Classification is dominated by deterministic signals — real status codes and well-known error strings matched against captured stdout/stderr:
400 unknown action→ hallucinated action vocabulary- Observed
404/405on the asserted route → the route doesn’t exist - Received-side
4xxcontract errors (assert 404 == 200) → the payload schema didn’t match reality
The received-vs-expected distinction is deliberate: assert 200 == 404 (expected 200, got 404) is a hallucination signal, but a negative test asserting a 4xx that got a 2xx stays a real finding. An LLM judge exists only as an adjudication hook for genuinely ambiguous output — the common cases never consult a model, so classification is fast, cheap, and stable.
Where you see it
Section titled “Where you see it”Generated scripts carry a verification badge in the UI. Quarantined scripts are excluded from coverage claims and from gate conclusions — a hallucinated test can never silently pass (or fail) your PR checks. When a script is self-healed or edited, re-execution re-runs the gate and the badge updates.