The preprint introduces V-JEPA Policy, a world-action model that couples future visual-state prediction with action generation. According to the abstract, a prominent line of recent WAMs works by adapting video generators or image-editing models that were pretrained at scale, thereby inheriting predictive knowledge from those models.

V-JEPA Policy takes a different route, building on predictive visual latents. The abstract suggests this approach is aimed at creating effective world-action models without relying on the video-generation or image-editing paradigm. The paper is available on arXiv under identifier 2609.37250.

Since only this single source was reviewed, the summary reflects its claims alone; no independent comparisons or conflicting findings are included.