A fact, then a closed door
A software agent learns one crucial fact in session one. That session ends. New requests, corrections and plausible distractions arrive. By session three, the agent faces a task whose solution depends on the earlier evidence.
Finding the fact is only half the job. The agent must use it correctly, change the repository and satisfy an executable check. Memory can deliver context to the workbench. It cannot pick up the tools.
One story, then measurement
The interactive begins with one fact crossing three sessions. Controlled measurement then expands that journey into 60 memory challenges under three seeds, producing 180 aligned final-session checks for each condition.
The honest ending
The experiment did not distinguish the two reference conditions. It also did not establish that they are equivalent. A benchmark should make bluffing harder for the system under test—and for the people describing the result.