Pub-AI: AI in Science

World Models Radar

A significance-filtered record of what actually moved world models forward.

Every new paper is summarized automatically and labelled. Editor picks are the ones a human judged to matter.

Tracks

Editor picks

Cosmos World Foundation Model Platform for Physical AI

2025-01-07 · NVIDIA, :, Niket Agarwal et al. · arXiv:2501.03575

Cosmos is NVIDIA's platform for 'world foundation models' for Physical AI: a video curation pipeline, pre-trained diffusion and autoregressive world models, examples of post-training for downstream tasks, and video tokenizers, released open-source with open-weight models.

Video World ModelscodeEditor pick

GAIA-1: A Generative World Model for Autonomous Driving

2023-09-29 · Anthony Hu, Lloyd Russell, Hudson Yeo et al. · arXiv:2309.17080

GAIA-1 is a generative world model for driving that takes video, text and action inputs, casts world modelling as next-token prediction over discrete tokens, and decodes with a video diffusion model to produce realistic driving scenes with control over ego-vehicle behaviour and scene features.

DrivingEditor pick

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

2025-06-11 · Mido Assran, Adrien Bardes, David Fan et al. · arXiv:2506.09985

Pretrains an action-free JEPA video model on over 1 million hours of internet video, then post-trains a latent action-conditioned world model (V-JEPA 2-AC) on under 62 hours of unlabeled robot video and uses it to plan pick-and-place with image goals, zero-shot on Franka arms in two labs.

JEPAcodeEditor pick

Latest

Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling

2026-09-30 · Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · arXiv:2609.31657

The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).

JEPAauto-summary

Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?

2026-09-30 · Yutong Feng, Bowen Liao, See Kiong Ng et al. · arXiv:2609.31680

The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.

JEPAauto-summary

CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense

2026-09-30 · Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan et al. · arXiv:2609.31893

CyberWorld is a Dreamer-style world model for autonomous cyber defense that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of a defended network. Using CyberWheel, the graph-based variant reaches the deploy_then_stop control after 3.6k–15.8k environment steps (vs model-free PPO needing 2.3M–3.1M steps or failing within a 3.2M budget).

Drivingauto-summary

Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

2026-09-30 · Jiaming Zhang, Wu Yang, Shuai Tao et al. · arXiv:2609.31938

Cache-aware Conv3D lowering for embedded generative video VAEs: supported causal Conv3D calls are expressed as batched spatial Conv2D while preserving pretrained weights, causal-cache semantics, convolution parameters, bias placement, and output layout. On 64-GB Jetson AGX Orin in the Cosmos3-Edge image-to-video pipeline, the authors report ~7.32× VAE-decoder speedup and 2.21× end-to-end speedup at 25 frames.

Video World Modelsauto-summary

Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement

2026-09-30 · Hitoshi Iyatomi · arXiv:2609.32268

The authors study Self-Reconstruction Dynamics (SRD), produced by repeatedly applying a frozen autoencoder to its own reconstruction, which forms transient image/latent trajectories. Although repeated self-reconstruction degrades fidelity, SRD encodes sample-specific correction information. They propose SRD-guided Reconstruction Refinement (SRD-RR), predicting a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. Across six datasets, SRD-RR recovers 38.6% (one) to 45.3% (two) of the empirically recoverable MSE gap.

Latent Dynamicsauto-summary

See all 126 papers → · Editor picks only