Tuesday, 22 September 2026

Search
Latent Digest

TECHNOLOGY, TRACKED ACROSS DISCIPLINES

Research Digest

Open-Vocabulary 3D Perception Extends Beyond Fixed Category Sets

Three papers push 3D object detection and segmentation beyond closed category lists, but they target different tasks and sensor inputs.

· 1 min read · 3 sources

All three papers start from the same limitation: a 3D detector or segmenter trained on a fixed set of object classes treats everything outside that set as invisible. Each work proposes an open-vocabulary alternative, but the papers diverge in task and input.

The first submission applies promptable segmentation to 3D object detection for autonomous driving, allowing queries for categories not present in the training boxes. SenseFuse targets 3D instance segmentation and combines image and shape encoders without needing labels, aiming to support robotic manipulation and spatial reasoning. The third paper restricts input to depth-only data for semantic segmentation, explicitly motivated by privacy preservation in indoor robots.

Agreement is at the level of goal: fixed category lists are too brittle for real-world perception. The difference is in constraints. Driving detectors can use rich visual cues, while privacy-aware indoor systems may be limited to depth. Together, the three works show open-vocabulary 3D perception is being specialized by application, not just unified by a single method.

Sources · 3

  1. 01Open-vocabulary 3D object detection with promptable segmentationarXiv
  2. 02SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance SegmentationarXiv
  3. 03Depth-Only Open-Vocabulary 3D Semantic Segmentation For Privacy-Preserving Robotic ApplicationsarXiv

More in Research Digest