Skip to content

Reuse & rediscovery

Memory is only useful if repetition is detectable. QualityMax surfaces reuse at three levels: the session, the test case, and the individual learning.

Every session’s sequence of tool calls is reduced to a stable key — effectively, a fingerprint of “what did this session actually do, in order.” Sessions sharing a key are grouped, the highest-scoring one becomes the representative, and the rest are reported as a normalized_negotiation_sequences group with an occurrences count. A group with occurrences: 4 means the same negotiation path was independently re-run four times — a direct signal that an agent (or four different agents) repeated work memory recall should have avoided.

Session scoring itself weighs verified success first, then completion, tool-call reliability, retained artifacts, and recency — a session with failed tool calls is explicitly penalized rather than scored the same as a clean run.

Case-level: what’s proven versus what’s fragile

Section titled “Case-level: what’s proven versus what’s fragile”

For a given test case, traversal walks its verified scripts and retained evidence to surface three things an agent should check before rediscovering them from scratch:

  • Stable locators — selectors that have held across at least two passing runs, ranked by how many passing runs used them
  • Fragile steps — steps previously flagged low-confidence or fragile, so a repair doesn’t have to re-learn that a given step is prone to breaking
  • Repeated failure signals — a fingerprint and rough classification (locator, timeout, assertion, navigation) for scripts that have failed the same way more than once

The design intent is explicit in the tool’s own guidance text: “Do not rediscover locators that held across passing runs.”

Learning-level: consolidation instead of accumulation

Section titled “Learning-level: consolidation instead of accumulation”

Individual learnings persist with support_count and contradiction_count, and near-duplicates are merged rather than left to pile up — see grounded learnings & evidence for how merging and the contradiction rule work.

Reuse detection is not just structural — it’s measured. A caller can report downstream_reuse back to get_project_memory after acting on a recalled memory, and whether a session called get_project_memory at all rolls up into the memory_reuse_rate metric on the memory graph. The goal is a system where “did memory actually get reused” is a number you can look at, not an assumption baked into the design.