Reuse & rediscovery
Memory is only useful if repetition is detectable. QualityMax surfaces reuse at three levels: the session, the test case, and the individual learning.
Session-level: duplicate negotiations
Section titled “Session-level: duplicate negotiations”Every session’s sequence of tool calls is reduced to a stable key — effectively, a fingerprint of “what did this session actually do, in order.” Sessions sharing a key are grouped, the highest-scoring one becomes the representative, and the rest are reported as a normalized_negotiation_sequences group with an occurrences count. A group with occurrences: 4 means the same negotiation path was independently re-run four times — a direct signal that an agent (or four different agents) repeated work memory recall should have avoided.
Session scoring itself weighs verified success first, then completion, tool-call reliability, retained artifacts, and recency — a session with failed tool calls is explicitly penalized rather than scored the same as a clean run.
Case-level: what’s proven versus what’s fragile
Section titled “Case-level: what’s proven versus what’s fragile”For a given test case, traversal walks its verified scripts and retained evidence to surface three things an agent should check before rediscovering them from scratch:
- Stable locators — selectors that have held across at least two passing runs, ranked by how many passing runs used them
- Fragile steps — steps previously flagged low-confidence or fragile, so a repair doesn’t have to re-learn that a given step is prone to breaking
- Repeated failure signals — a fingerprint and rough classification (locator, timeout, assertion, navigation) for scripts that have failed the same way more than once
The design intent is explicit in the tool’s own guidance text: “Do not rediscover locators that held across passing runs.”
Learning-level: consolidation instead of accumulation
Section titled “Learning-level: consolidation instead of accumulation”Individual learnings persist with support_count and contradiction_count, and near-duplicates are merged rather than left to pile up — see grounded learnings & evidence for how merging and the contradiction rule work.
Closing the loop
Section titled “Closing the loop”Reuse detection is not just structural — it’s measured. A caller can report downstream_reuse back to get_project_memory after acting on a recalled memory, and whether a session called get_project_memory at all rolls up into the memory_reuse_rate metric on the memory graph. The goal is a system where “did memory actually get reused” is a number you can look at, not an assumption baked into the design.