Panoramic images offer a complete 360-degree field of view, which is valuable for embodied perception. However, different robotic platforms—such as vehicles, drones, wearables, and quadrupeds—have very different observation viewpoints and spatial layouts. The authors of this paper argue that these "cross-embodiment observation shifts" make consistent and reliable panoramic perception difficult, and they note that systematic studies of this problem remain limited.

To fill that gap, the authors propose a new task called Cross-Embodiment Open Panoramic Segmentation and introduce EmbPASS, a multi-platform benchmark spanning Vehicle, Drone, Wearable, and Quadruped platforms under a unified semantic taxonomy. This provides a common testbed for studying how well segmentation models transfer across heterogeneous embodiments.

They also present EPONet, an open-vocabulary panoramic semantic segmentation network. It uses two components: a Relation-Aware Metric Adapter (RAMA) for enhanced spatial modeling and a Content-Adaptive Semantic Transfer (CAST) module for semantic transfer under varied embodied observations. In experiments, EPONet achieves the best platform-balanced performance on EmbPASS with 35.82% mean IoU, outperforming the strongest baseline by 1.10%, while remaining competitive on existing panoramic segmentation benchmarks.

The authors state that the source code and the EmbPASS benchmark will be made publicly available, which could support further research on cross-embodiment perception in robotics and computer vision.