MemoryAgentBench's Conflict Resolution split is often read as a test of whether agents selectively forget outdated facts. A new arXiv paper questions that interpretation by executing the benchmark's own rule: the newest statement about a fact wins. The authors implement this as a zero-learning resolver, frozen on a single fact condition, and compare its scores against the benchmark's published conflict-resolution results.
The paper's approach is deliberately simple. Instead of training or updating an agent, the resolver just applies the benchmark's scoring rule mechanically. The finding is that this frozen last-write-wins baseline can reproduce the conflict-resolution scores, suggesting that the benchmark may be measuring straightforward bookkeeping rather than memory dynamics.
The authors caution that the Conflict Resolution split may not be a valid measure of selective forgetting. Their analysis is a reminder that benchmark scores need to be read against simple baselines before attributing them to complex cognitive abilities. As the only source here, the paper presents a single, focused argument; no independent replication or counter-evidence is included.