Language-conditioned robot policies have become a key research area, but most current methods use a single, high-level task description to guide learning. That coarse conditioning may miss important details about how to execute sub-steps, especially in complex tasks.
The new paper introduces a multi-granularity framework that explicitly decomposes instructions into finer components. Instead of feeding the whole instruction at once, the policy learns from sub-instructions at different levels of detail, allowing it to align actions more precisely with each stage of the task.
The authors argue that this decomposition helps imitation learning by providing richer supervision signals. While the abstract does not report quantitative results, the approach suggests a promising direction for making language-guided policies more responsive to the structure of real-world manipulation tasks.