Skip to content

Hallucination gate

An AI-generated test can look perfect and test nothing. The hallucination gate is the pipeline stage that tells the difference — by executing the generated test against the live target and classifying the result.

The founding incident: a dogfood crawl against a real product generated 40 test scripts, and 38 of them asserted against an API contract that didn’t exist — every create call returned 400 unknown action. Every test “passed review” as code. All of them were hallucinations. Static inspection cannot catch this class of failure; only running against reality can.

State Meaning What happens
Verified Ran green against the live target Counts as automation
Failed live Target exists, behaves wrong A real found bug — keep it. A failing test that reproduces a genuine defect is a legitimate test
Hallucinated Asserted endpoint/contract/selector doesn’t exist Quarantined — never counts as the test case’s automation
Inconclusive Infra/transient failure (timeout, 502–504, connection refused) Badge only, no verdict

(A fifth state, unverified, marks scripts never executed — e.g. syntax-only saves.)

The distinction that matters: a hallucination is the only state that must never count as coverage. failed_live is gold — an automated reproduction of a real bug. hallucinated is debt wearing a test’s clothes.

Classification is dominated by deterministic signals — real status codes and well-known error strings matched against captured stdout/stderr:

  • 400 unknown action → hallucinated action vocabulary
  • Observed 404/405 on the asserted route → the route doesn’t exist
  • Received-side 4xx contract errors (assert 404 == 200) → the payload schema didn’t match reality

The received-vs-expected distinction is deliberate: assert 200 == 404 (expected 200, got 404) is a hallucination signal, but a negative test asserting a 4xx that got a 2xx stays a real finding. An LLM judge exists only as an adjudication hook for genuinely ambiguous output — the common cases never consult a model, so classification is fast, cheap, and stable.

Generated scripts carry a verification badge in the UI. Quarantined scripts are excluded from coverage claims and from gate conclusions — a hallucinated test can never silently pass (or fail) your PR checks. When a script is self-healed or edited, re-execution re-runs the gate and the badge updates.