Skip to content

Example: repairing with vs. without memory

Take a common task: a Playwright test broke after a UI change, and an agent needs to repair it.

Without recalling project memory first, the agent has to rebuild context from primitives: list the project’s test cases to find the right one, fetch the test case, fetch the current script, run it to see the actual failure, then reason from the raw error about which locator moved — with no signal about whether that locator was already known to be fragile, or whether a different locator in the same case has already failed the same way twice before.

Calling get_project_memory with intent="repair" and a query describing the task returns, in one call: the case’s stable locators (proven across multiple passing runs, so the agent knows what not to touch), any steps already flagged fragile from prior evidence, and repeated-failure signals with a classification (locator, timeout, assertion, navigation) if this script has failed this way before. The agent then calls get_script for exactly the one current, verified script — not the whole project’s script list — and repairs with the fragile step already identified instead of rediscovering it from a stack trace.

A measured production MCP session — 75 tool calls, no memory recall — came in at roughly 70,600 estimated payload tokens. Two tool families, update_script and run_tests, accounted for about half of that; adding get_script and get_test_case pushed the concentration to roughly three-quarters. That matches get_project_memory’s own guidance to agents almost exactly: “call get_script only for the most relevant verified script to keep context and cost low.” Recall doesn’t eliminate those calls — it aims them, so a repair fetches the one script it needs instead of re-deriving which one that is.

This isn’t presented as a fixed savings percentage, because the actual reduction depends on how much prior verified history a project has. What’s measurable instead: whether a session called get_project_memory before acting (rolled up as memory_reuse_rate on the memory graph), and whether a caller reports back that a recalled memory was actually used (downstream_reuse). Reuse is tracked as a metric the team is actively driving up, not claimed as a number already banked.