Tuesday, 22 September 2026

Search
Latent Digest

TECHNOLOGY, TRACKED ACROSS DISCIPLINES

Research Digest

Skeleton-guided world models turn human videos into robot training data

Seven arXiv papers demonstrate how human demonstrations, skeletons, and world models can reduce reliance on costly robot-collected training data.

· 2 min read · 7 sources

Robot demonstrations are expensive and often cover a narrow slice of task variations, argue the authors of Skel-WAM, KnowDemo, and HuRo. They each treat human videos as a cheap, diverse substitute or complement. The abstracts agree on the goal but diverge in method: one conditions on hand skeletons, another injects knowledge into demonstration generation, and a third "robotizes" videos for vision-language-action pretraining.

Two papers explicitly build on world-action models. The MT-WAM abstract cites Fast-WAM's finding that video-action co-training helps control without future-video generation, and proposes reorienting the one-pass predictive representation toward action generation. The Skel-WAM papers go further by using skeleton guidance to carry manipulation experience across embodiments, with one claiming zero-shot cross-embodiment manipulation despite changes in appearance and action dimensionality.

On humanoid systems, the focus shifts from arm and hand control to whole-body skills. One paper trains scene-aware locomotion through 3D clutter from immersive human demonstrations, addressing an environment type the authors say is underexplored. Another learns distance-conditioned object transport from a single retargeted motion clip, aiming to generalize beyond the exact transport outcome in the reference clip.

Taken together, the seven preprints show a research cluster around squeezing reusable training signal from human motion. The differences are mostly in where the generalization comes from—skeleton abstractions, knowledge priors, distance conditioning, or scene awareness—rather than in the fundamental premise that human data can reduce the cost of robot learning.

Sources · 7

  1. 01KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human VideosarXiv
  2. 02MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action GenerationarXiv
  3. 03HuRo: Robotizing Human Videos for Scalable VLA PretrainingarXiv
  4. 04Learning Scene-Aware Humanoid Locomotion through 3D Clutter from Immersive Human DemonstrationsarXiv
  5. 05Learning Distance-Conditioned Object Transport for Humanoid Loco-Manipulation from a Single Motion CliparXiv
  6. 06SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment ManipulationarXiv
  7. 07Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation TransferarXiv

More in Research Digest