A new paper on arXiv, titled "PreferenceFlow: Test-Time Guidance of Flow-Matching Robot Policies from Human Interventions," addresses a common failure mode for learned robot policies. The authors note that flow-matching policies, while capable of representing complex behaviors, often make local errors when deployed under distribution shift—that is, when the test environment differs from training.

The abstract points out that many reinforcement learning methods for policy improvement rely on reward signals. Such signals are frequently difficult to obtain or specify in practice. The paper's proposal, PreferenceFlow, appears to sidestep this by instead using human interventions to guide the policy at test time, rather than requiring explicit rewards or full retraining.

Because only the abstract is available, the details of the mechanism are not yet public. The work is announced as a new arXiv preprint, and the full methodology will presumably be described in the forthcoming paper.