A significance-filtered record of what actually moved world models forward.
Every new paper is summarized automatically and labelled. Editor picks are the ones a human judged to matter.
Editor picks
2025-01-07 · NVIDIA, :, Niket Agarwal et al. · arXiv:2501.03575
Cosmos is NVIDIA's platform for 'world foundation models' for Physical AI: a video curation pipeline, pre-trained diffusion and autoregressive world models, examples of post-training for downstream tasks, and video tokenizers, released open-source with open-weight models.
Video World ModelscodeEditor pick
2023-09-29 · Anthony Hu, Lloyd Russell, Hudson Yeo et al. · arXiv:2309.17080
GAIA-1 is a generative world model for driving that takes video, text and action inputs, casts world modelling as next-token prediction over discrete tokens, and decodes with a video diffusion model to produce realistic driving scenes with control over ego-vehicle behaviour and scene features.
DrivingEditor pick
2025-06-11 · Mido Assran, Adrien Bardes, David Fan et al. · arXiv:2506.09985
Pretrains an action-free JEPA video model on over 1 million hours of internet video, then post-trains a latent action-conditioned world model (V-JEPA 2-AC) on under 62 hours of unlabeled robot video and uses it to plan pick-and-place with image goals, zero-shot on Franka arms in two labs.
JEPAcodeEditor pick
Latest
2026-09-30 · Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · arXiv:2609.31657
The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).
JEPAauto-summary
2026-09-30 · Yutong Feng, Bowen Liao, See Kiong Ng et al. · arXiv:2609.31680
The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.
JEPAauto-summary
2026-09-30 · Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan et al. · arXiv:2609.31893
CyberWorld is a Dreamer-style world model for autonomous cyber defense that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of a defended network. Using CyberWheel, the graph-based variant reaches the deploy_then_stop control after 3.6k–15.8k environment steps (vs model-free PPO needing 2.3M–3.1M steps or failing within a 3.2M budget).
Drivingauto-summary
2026-09-30 · Jiaming Zhang, Wu Yang, Shuai Tao et al. · arXiv:2609.31938
Cache-aware Conv3D lowering for embedded generative video VAEs: supported causal Conv3D calls are expressed as batched spatial Conv2D while preserving pretrained weights, causal-cache semantics, convolution parameters, bias placement, and output layout. On 64-GB Jetson AGX Orin in the Cosmos3-Edge image-to-video pipeline, the authors report ~7.32× VAE-decoder speedup and 2.21× end-to-end speedup at 25 frames.
Video World Modelsauto-summary
2026-09-30 · Hitoshi Iyatomi · arXiv:2609.32268
The authors study Self-Reconstruction Dynamics (SRD), produced by repeatedly applying a frozen autoencoder to its own reconstruction, which forms transient image/latent trajectories. Although repeated self-reconstruction degrades fidelity, SRD encodes sample-specific correction information. They propose SRD-guided Reconstruction Refinement (SRD-RR), predicting a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. Across six datasets, SRD-RR recovers 38.6% (one) to 45.3% (two) of the empirically recoverable MSE gap.
Latent Dynamicsauto-summary
See all 126 papers → · Editor picks only