Skip to content

Grounded learnings & evidence

Grounding is enforced in three separate places, each catching a different way ungrounded content can creep in.

Before generating a new test, QualityMax retrieves verified, passing tests from the same project by embedding similarity and injects the closest matches as few-shot examples — each one labeled with its similarity score, for example Example passing playwright test (similarity 0.47). The examples are wrapped as untrusted retrieved data, not instructions, with the same control-character and role-spoofing defenses used everywhere recalled text reaches a model.

This retrieval degrades safely: a project with no prior passing tests, an embedding failure, or a zero-result search all fall back to generation without examples rather than blocking the request. Every attempt is logged with a specific reason — no_query, no_project_id, embed_failed, zero_results, ok, and so on — so an empty grounding block from a brand-new project is distinguishable from a broken retrieval path.

Free-form project_learnings text is easy to write and easy for a fact to go stale in. Before it’s surfaced back to an agent as grounded_project_learnings, each line is checked for a verifiable anchor — a URL, an element ID, a data-* attribute, or a getBy*/locator() call — and that anchor must actually appear in the project’s verified evidence (passing scripts and retained evidence packages). A line with no anchor, or an anchor that doesn’t match anything verified, is rejected with a stated reason rather than silently dropped; the full set of accepted and rejected lines is available, along with a 0–100 gate score.

For repository-derived work, evidence starts even earlier: a deterministic scanner reads the repository’s routes, models, functions, validation rules, existing tests, and README claims into a stable evidence catalog before any model sees the code. The rule that follows from that is strict — the model is a planner, not a source of truth. A test area or test case that cites an evidence ID the scanner never produced is rejected before it can be persisted. get_discovery_graph exposes the result as a {nodes, edges, summary} graph: nodes are evidence, test areas, generated cases, and scripts; edges are contains, supports, produces, and automates; and the summary reports a grounding rate — the share of test cases that actually trace back to cited evidence versus ones carried over from before the evidence catalog existed (legacy_unverified).

Learnings aren’t append-only. Near-duplicate learnings within the same project and type are found by a combination of vector similarity and exact normalized-text matching, then merged into one survivor that accumulates the group’s support and contradiction counts. A contradicted learning can never become the survivor of a merge — evidence that something used to be true doesn’t let it outvote evidence that it no longer is.

Every layer follows the same rule: a generation, a learning, or a recalled fact is only as trustworthy as the citation it carries, and the citation is checked against something that was actually observed — a scanner pass, a passing test run, or a prior verified session — not asserted by the model that produced it.