Skeleton-guided world models turn human videos into robot training data
Seven arXiv papers demonstrate how human demonstrations, skeletons, and world models can reduce reliance on costly robot-collected training data.
Robot demonstrations are expensive and often cover a narrow slice of task variations, argue the authors of Skel-WAM, KnowDemo, and HuRo. They each treat human videos as a cheap, diverse substitute or complement. The abstracts agree on the goal but diverge in method: one conditions on hand skeletons, another injects knowledge into demonstration generation, and a third "robotizes" videos for vision-language-action pretraining.
Two papers explicitly build on world-action models. The MT-WAM abstract cites Fast-WAM's finding that video-action co-training helps control without future-video generation, and proposes reorienting the one-pass predictive representation toward action generation. The Skel-WAM papers go further by using skeleton guidance to carry manipulation experience across embodiments, with one claiming zero-shot cross-embodiment manipulation despite changes in appearance and action dimensionality.
On humanoid systems, the focus shifts from arm and hand control to whole-body skills. One paper trains scene-aware locomotion through 3D clutter from immersive human demonstrations, addressing an environment type the authors say is underexplored. Another learns distance-conditioned object transport from a single retargeted motion clip, aiming to generalize beyond the exact transport outcome in the reference clip.
Taken together, the seven preprints show a research cluster around squeezing reusable training signal from human motion. The differences are mostly in where the generalization comes from—skeleton abstractions, knowledge priors, distance conditioning, or scene awareness—rather than in the fundamental premise that human data can reduce the cost of robot learning.
Sources · 7
- KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos
- MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining
- Learning Scene-Aware Humanoid Locomotion through 3D Clutter from Immersive Human Demonstrations
- Learning Distance-Conditioned Object Transport for Humanoid Loco-Manipulation from a Single Motion Clip
- SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation
- Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer
More in Research Digest
Can LLM Agents Design Chips From Higher-Level Abstractions?
A new preprint asks whether large language model agents can outperform RTL-level approaches by designing chips from higher-level abstractions.
Research Digest: Memory and Cooperation in Multi-Agent Vision
New papers explore how vision-language agents can share memory and arbitrate roles, while other work tackles compact representations and multi-channel imaging.
New Papers Probe the Hidden Costs and Risks of LLM Reasoning Traces
Six recent arXiv papers examine what happens inside chain-of-thought reasoning, showing that intermediate traces can be a liability as much as a capability.
New AI Research Spans Networks, Economy, Art, and Tools
Five independent papers highlight AI's expanding footprint from network optimization to cultural critique.