Open-Vocabulary 3D Perception Extends Beyond Fixed Category Sets
Three papers push 3D object detection and segmentation beyond closed category lists, but they target different tasks and sensor inputs.
All three papers start from the same limitation: a 3D detector or segmenter trained on a fixed set of object classes treats everything outside that set as invisible. Each work proposes an open-vocabulary alternative, but the papers diverge in task and input.
The first submission applies promptable segmentation to 3D object detection for autonomous driving, allowing queries for categories not present in the training boxes. SenseFuse targets 3D instance segmentation and combines image and shape encoders without needing labels, aiming to support robotic manipulation and spatial reasoning. The third paper restricts input to depth-only data for semantic segmentation, explicitly motivated by privacy preservation in indoor robots.
Agreement is at the level of goal: fixed category lists are too brittle for real-world perception. The difference is in constraints. Driving detectors can use rich visual cues, while privacy-aware indoor systems may be limited to depth. Together, the three works show open-vocabulary 3D perception is being specialized by application, not just unified by a single method.
Sources · 3
More in Research Digest
Can LLM Agents Design Chips From Higher-Level Abstractions?
A new preprint asks whether large language model agents can outperform RTL-level approaches by designing chips from higher-level abstractions.
Research Digest: Memory and Cooperation in Multi-Agent Vision
New papers explore how vision-language agents can share memory and arbitrate roles, while other work tackles compact representations and multi-channel imaging.
New Papers Probe the Hidden Costs and Risks of LLM Reasoning Traces
Six recent arXiv papers examine what happens inside chain-of-thought reasoning, showing that intermediate traces can be a liability as much as a capability.
New AI Research Spans Networks, Economy, Art, and Tools
Five independent papers highlight AI's expanding footprint from network optimization to cultural critique.