Digest
3 papers
2025-06-11 · Mido Assran, Adrien Bardes, David Fan et al. · arXiv:2506.09985
Pretrains an action-free JEPA video model on over 1 million hours of internet video, then post-trains a latent action-conditioned world model (V-JEPA 2-AC) on under 62 hours of unlabeled robot video and uses it to plan pick-and-place with image goals, zero-shot on Franka arms in two labs.
JEPAcodeEditor pick
2025-01-07 · NVIDIA, :, Niket Agarwal et al. · arXiv:2501.03575
Cosmos is NVIDIA's platform for 'world foundation models' for Physical AI: a video curation pipeline, pre-trained diffusion and autoregressive world models, examples of post-training for downstream tasks, and video tokenizers, released open-source with open-weight models.
Video World ModelscodeEditor pick
2023-09-29 · Anthony Hu, Lloyd Russell, Hudson Yeo et al. · arXiv:2309.17080
GAIA-1 is a generative world model for driving that takes video, text and action inputs, casts world modelling as next-token prediction over discrete tokens, and decodes with a video diffusion model to produce realistic driving scenes with control over ego-vehicle behaviour and scene features.
DrivingEditor pick