Example: repairing with vs. without memory
Take a common task: a Playwright test broke after a UI change, and an agent needs to repair it.
Before: no memory recall
Section titled “Before: no memory recall”Without recalling project memory first, the agent has to rebuild context from primitives: list the project’s test cases to find the right one, fetch the test case, fetch the current script, run it to see the actual failure, then reason from the raw error about which locator moved — with no signal about whether that locator was already known to be fragile, or whether a different locator in the same case has already failed the same way twice before.
After: recall first
Section titled “After: recall first”Calling get_project_memory with intent="repair" and a query describing the task returns, in one call: the case’s stable locators (proven across multiple passing runs, so the agent knows what not to touch), any steps already flagged fragile from prior evidence, and repeated-failure signals with a classification (locator, timeout, assertion, navigation) if this script has failed this way before. The agent then calls get_script for exactly the one current, verified script — not the whole project’s script list — and repairs with the fragile step already identified instead of rediscovering it from a stack trace.
What actually drives the cost this avoids
Section titled “What actually drives the cost this avoids”A measured production MCP session — 75 tool calls, no memory recall — came in at roughly 70,600 estimated payload tokens. Two tool families, update_script and run_tests, accounted for about half of that; adding get_script and get_test_case pushed the concentration to roughly three-quarters. That matches get_project_memory’s own guidance to agents almost exactly: “call get_script only for the most relevant verified script to keep context and cost low.” Recall doesn’t eliminate those calls — it aims them, so a repair fetches the one script it needs instead of re-deriving which one that is.
Measuring it, not assuming it
Section titled “Measuring it, not assuming it”This isn’t presented as a fixed savings percentage, because the actual reduction depends on how much prior verified history a project has. What’s measurable instead: whether a session called get_project_memory before acting (rolled up as memory_reuse_rate on the memory graph), and whether a caller reports back that a recalled memory was actually used (downstream_reuse). Reuse is tracked as a metric the team is actively driving up, not claimed as a number already banked.