An agent that remembers what it did on a web page has to decide when two pages are the same. Memories built on observation similarity merge pages that look alike but behave differently, and GUIs are full of such pages—two tabs that look alike can lead to very different states. A memory that cannot tell them apart will mislead the agent later.
The paper proposes action-conditioned bisimulation for building GUI agent memory. Rather than grouping pages by visual similarity, the idea is to group them by how they respond to actions. Two pages count as the same only when the agent's actions on them lead to equivalent outcomes. This shifts the criterion for sameness from appearance to behavior.
The approach directly targets a known failure mode of observation-based memory. For an agent that acts on a page, what matters is not whether two pages look alike but whether they afford the same next steps. Conditioning memory on actions makes the resulting representation line up with the agent's own experience.