2026-09-30 · Yu Xu, Yuxin Zhang, Xiao Yang et al. · arXiv:2609.38140
The authors argue that token-wise load-balanced MoE routing forms a “uniformity trap” for video, scattering coherent spatiotemporal tokens across experts and causing fragmentation and distortion. They propose SplitMoE, splitting experts into semantic and generic roles with prototype-guided semantic routing and pull-push regularization, reporting faster convergence and improved routing coherence and video generation quality under matched activated-parameter budgets.
LongLive-Plug introduces a once-for-all distillation framework that learns reusable LoRA adapters on a base video diffusion model. The adapters cover single-pass classifier-free guidance, few-step (four-step) sampling, and long-context error correction for causal autoregressive generation, enabling training-free plug-and-play deployment across compatible downstream models.
2026-09-30 · Haoyi Jiang, Liu Liu, Xinjiang Wang et al. · arXiv:2609.38163
ReWAM is a representation-centric world-action model that uses frozen pre-trained DINO features and builds compact temporal world states via a Temporal Representation Bottleneck (TRB). Action-Grounded Representation Shaping (AGRS) routes only action-loss gradients to the TRB, letting the policy shape what the representation encodes. The authors report 93.6% success on RoboTwin 2.0 without generative video pre-training; on RoboDojo they report average score 9.42 and success rate 5.82% without embodied pre-training and 12.29 and 8.28% with ~600 hours.
Rho is an open-weights family of 5B-parameter vision-language-action (VLA) models for bimanual robotic manipulation across three embodiments (YAM Box, UR AI Trainer, FR3 Duo). The authors design regularized progressive adaptation with embodiment midtraining, and show improved downstream adaptation in simulation and physical-robot experiments. They also add online latent adaptation that learns from corrective feedback using as few as 15 corrected episodes.
2026-09-30 · Bingchen Yao, Haobo Xu, Haokun Lin et al. · arXiv:2609.38169
STEPQuant is a spatial-temporal post-training quantization method for Delta-rule linear-attention recurrent states. It allocates quantization precision using lifetime-aware bit allocation (memory error persistence) and key-row-aware dual-axis fitting (key-row impact and row/column magnitude structure). On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, it matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 at 4-bit, with large serving memory reductions.
PRISM uses grounded video-to-video (V2V) generation to create hundreds of counterfactual human–object interaction videos from four real seed videos. A contact-anchored real-to-sim pipeline reconstructs and retargets monocular videos into physically plausible robot–object trajectories. A single depth-based humanoid policy is trained and deployed on a real robot zero-shot (no real-world fine-tuning) using onboard depth for pick–carry–drop on diverse objects.
2025-06-11 · Mido Assran, Adrien Bardes, David Fan et al. · arXiv:2506.09985
Pretrains an action-free JEPA video model on over 1 million hours of internet video, then post-trains a latent action-conditioned world model (V-JEPA 2-AC) on under 62 hours of unlabeled robot video and uses it to plan pick-and-place with image goals, zero-shot on Franka arms in two labs.
2025-01-07 · NVIDIA, :, Niket Agarwal et al. · arXiv:2501.03575
Cosmos is NVIDIA's platform for 'world foundation models' for Physical AI: a video curation pipeline, pre-trained diffusion and autoregressive world models, examples of post-training for downstream tasks, and video tokenizers, released open-source with open-weight models.
2024-05-27 · Shenyuan Gao, Jiazhi Yang, Li Chen et al. · arXiv:2405.17398
Vista is a generalizable driving world model that predicts at high resolution, supports many action controls (command, goal point, trajectory, angle and speed), and can act as a reward function to evaluate actions without ground-truth actions.
2024-05-20 · Eloi Alonso, Adam Jelley, Vincent Micheli et al. · arXiv:2405.12399
DIAMOND trains an RL agent inside a diffusion world model, arguing that discrete-latent compression can lose visual details that matter for control; it reaches a mean human-normalized score of 1.46 on Atari 100k.
2024-02-23 · Jake Bruce, Michael Dennis, Ashley Edwards et al. · arXiv:2402.15391
Genie is an 11B-parameter generative interactive environment trained without action labels on unlabelled internet videos; a spatiotemporal video tokenizer, autoregressive dynamics model and latent action model let users act frame by frame in generated worlds.
2024-02-15 · Adrien Bardes, Quentin Garrido, Jean Ponce et al. · arXiv:2404.08471
V-JEPA trains video models solely with a feature-prediction objective (predicting masked spatio-temporal regions in representation space), with no pretrained image encoders, text, negatives or reconstruction, on 2 million public videos.
2023-10-25 · Nicklas Hansen, Hao Su, Xiaolong Wang · arXiv:2310.16828
TD-MPC2 improves TD-MPC, which does local trajectory optimization in the latent space of a decoder-free (implicit) world model; one set of hyperparameters works across 104 online RL tasks, and a single 317M-parameter agent performs 80 tasks across domains and embodiments.
2023-10-09 · Sherry Yang, Yilun Du, Kamyar Ghasemipour et al. · arXiv:2310.06114
UniSim learns a universal simulator of real-world interaction by orchestrating diverse datasets (image, robotics, navigation) so it can simulate the visual outcome of both high-level instructions and low-level controls; policies trained purely inside it deploy zero-shot on real robots.
2023-09-29 · Anthony Hu, Lloyd Russell, Hudson Yeo et al. · arXiv:2309.17080
GAIA-1 is a generative world model for driving that takes video, text and action inputs, casts world modelling as next-token prediction over discrete tokens, and decodes with a video diffusion model to produce realistic driving scenes with control over ego-vehicle behaviour and scene features.
DriveDreamer is a diffusion-based world model learned from real-world driving data (evaluated on nuScenes) that generates controllable driving videos and future driving policies, trained in two stages (structured traffic constraints first, then future-state prediction).
I-JEPA learns image representations by predicting the representations of several target blocks of an image from a single context block, in representation space and with no hand-crafted data augmentations.
2023-01-10 · Danijar Hafner, Jurgis Pasukonis, Jimmy Ba et al. · arXiv:2301.04104
DreamerV3 is a general model-based RL algorithm that learns a world model and improves behavior by imagining future scenarios; with a single configuration it outperforms specialized methods across over 150 tasks, and is the first to collect diamonds in Minecraft from scratch without human data.
2022-09-01 · Vincent Micheli, Eloi Alonso, François Fleuret · arXiv:2209.00588
IRIS is a data-efficient RL agent that learns inside a world model built from a discrete autoencoder plus an autoregressive Transformer, and reaches a mean human-normalized score of 1.046 on Atari 100k (about two hours of gameplay).
2022-06-28 · Philipp Wu, Alejandro Escontrela, Danijar Hafner et al. · arXiv:2206.14176
DayDreamer applies the Dreamer algorithm to four physical robots that learn online in the real world without simulators, using the same hyperparameters across tasks: a quadruped, two robot arms and a wheeled robot.