Skip to content

05 / 12

Can a fact survive the distance?

Independent research · Agent memory

DreamBench-SWE tests whether a software agent can carry necessary evidence across sessions and still complete the software work that decides success.

Open project (opens in a new tab)

A fact, then a closed door

A software agent learns one crucial fact in session one. That session ends. New requests, corrections and plausible distractions arrive. By session three, the agent faces a task whose solution depends on the earlier evidence.

Finding the fact is only half the job. The agent must use it correctly, change the repository and satisfy an executable check. Memory can deliver context to the workbench. It cannot pick up the tools.

One story, then measurement

The interactive begins with one fact crossing three sessions. Controlled measurement then expands that journey into 60 memory challenges under three seeds, producing 180 aligned final-session checks for each condition.

The honest ending

The experiment did not distinguish the two reference conditions. It also did not establish that they are equivalent. A benchmark should make bluffing harder for the system under test—and for the people describing the result.