Pub-AI: AI in Science

Digest

Reset

126 papers

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

2026-09-30 · Yu Xu, Yuxin Zhang, Xiao Yang et al. · arXiv:2609.38140

The authors argue that token-wise load-balanced MoE routing forms a “uniformity trap” for video, scattering coherent spatiotemporal tokens across experts and causing fragmentation and distortion. They propose SplitMoE, splitting experts into semantic and generic roles with prototype-guided semantic routing and pull-push regularization, reporting faster convergence and improved routing coherence and video generation quality under matched activated-parameter budgets.

Video World Modelsauto-summary

LongLive-Plug: Once-for-All Distillation for Video Generation

2026-09-30 · Shuai Yang, Luozhou Wang, Wei Huang et al. · arXiv:2609.38154

LongLive-Plug introduces a once-for-all distillation framework that learns reusable LoRA adapters on a base video diffusion model. The adapters cover single-pass classifier-free guidance, few-step (four-step) sampling, and long-context error correction for causal autoregressive generation, enabling training-free plug-and-play deployment across compatible downstream models.

Video World Modelsauto-summary

Rethinking Representations for World-Action Modeling

2026-09-30 · Haoyi Jiang, Liu Liu, Xinjiang Wang et al. · arXiv:2609.38163

ReWAM is a representation-centric world-action model that uses frozen pre-trained DINO features and builds compact temporal world states via a Temporal Representation Bottleneck (TRB). Action-Grounded Representation Shaping (AGRS) routes only action-loss gradients to the TRB, letting the policy shape what the representation encodes. The authors report 93.6% success on RoboTwin 2.0 without generative video pre-training; on RoboDojo they report average score 9.42 and success rate 5.82% without embodied pre-training and 12.29 and 8.28% with ~600 hours.

Drivingcodeauto-summary

Rho: A Foundation for Efficiently Adaptable VLA Models

2026-09-30 · Rho Team, Simran Bagaria, Daphne Chen et al. · arXiv:2609.38164

Rho is an open-weights family of 5B-parameter vision-language-action (VLA) models for bimanual robotic manipulation across three embodiments (YAM Box, UR AI Trainer, FR3 Duo). The authors design regularized progressive adaptation with embodiment midtraining, and show improved downstream adaptation in simulation and physical-robot experiments. They also add online latent adaptation that learns from corrective feedback using as few as 15 corrected episodes.

Drivingauto-summary

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

2026-09-30 · Bingchen Yao, Haobo Xu, Haokun Lin et al. · arXiv:2609.38169

STEPQuant is a spatial-temporal post-training quantization method for Delta-rule linear-attention recurrent states. It allocates quantization precision using lifetime-aware bit allocation (memory error persistence) and key-row-aware dual-axis fitting (key-row impact and row/column magnitude structure). On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, it matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 at 4-bit, with large serving memory reductions.

Drivingcodeauto-summary

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

2026-09-30 · Zihan Wang, Zhen Wu, Pieter Abbeel et al. · arXiv:2609.38172

PRISM uses grounded video-to-video (V2V) generation to create hundreds of counterfactual human–object interaction videos from four real seed videos. A contact-anchored real-to-sim pipeline reconstructs and retargets monocular videos into physically plausible robot–object trajectories. A single depth-based humanoid policy is trained and deployed on a real robot zero-shot (no real-world fine-tuning) using onboard depth for pick–carry–drop on diverse objects.

Drivingauto-summary

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

2025-06-11 · Mido Assran, Adrien Bardes, David Fan et al. · arXiv:2506.09985

Pretrains an action-free JEPA video model on over 1 million hours of internet video, then post-trains a latent action-conditioned world model (V-JEPA 2-AC) on under 62 hours of unlabeled robot video and uses it to plan pick-and-place with image goals, zero-shot on Franka arms in two labs.

JEPAcodeEditor pick

Cosmos World Foundation Model Platform for Physical AI

2025-01-07 · NVIDIA, :, Niket Agarwal et al. · arXiv:2501.03575

Cosmos is NVIDIA's platform for 'world foundation models' for Physical AI: a video curation pipeline, pre-trained diffusion and autoregressive world models, examples of post-training for downstream tasks, and video tokenizers, released open-source with open-weight models.

Video World ModelscodeEditor pick

Diffusion for World Modeling: Visual Details Matter in Atari

2024-05-20 · Eloi Alonso, Adam Jelley, Vincent Micheli et al. · arXiv:2405.12399

DIAMOND trains an RL agent inside a diffusion world model, arguing that discrete-latent compression can lose visual details that matter for control; it reaches a mean human-normalized score of 1.46 on Atari 100k.

Latent Dynamicsauto-summary

Genie: Generative Interactive Environments

2024-02-23 · Jake Bruce, Michael Dennis, Ashley Edwards et al. · arXiv:2402.15391

Genie is an 11B-parameter generative interactive environment trained without action labels on unlabelled internet videos; a spatiotemporal video tokenizer, autoregressive dynamics model and latent action model let users act frame by frame in generated worlds.

Video World Modelsauto-summary

Revisiting Feature Prediction for Learning Visual Representations from Video

2024-02-15 · Adrien Bardes, Quentin Garrido, Jean Ponce et al. · arXiv:2404.08471

V-JEPA trains video models solely with a feature-prediction objective (predicting masked spatio-temporal regions in representation space), with no pretrained image encoders, text, negatives or reconstruction, on 2 million public videos.

JEPAcodeauto-summary

TD-MPC2: Scalable, Robust World Models for Continuous Control

2023-10-25 · Nicklas Hansen, Hao Su, Xiaolong Wang · arXiv:2310.16828

TD-MPC2 improves TD-MPC, which does local trajectory optimization in the latent space of a decoder-free (implicit) world model; one set of hyperparameters works across 104 online RL tasks, and a single 317M-parameter agent performs 80 tasks across domains and embodiments.

Latent Dynamicsauto-summary

Learning Interactive Real-World Simulators

2023-10-09 · Sherry Yang, Yilun Du, Kamyar Ghasemipour et al. · arXiv:2310.06114

UniSim learns a universal simulator of real-world interaction by orchestrating diverse datasets (image, robotics, navigation) so it can simulate the visual outcome of both high-level instructions and low-level controls; policies trained purely inside it deploy zero-shot on real robots.

Video World Modelsauto-summary

GAIA-1: A Generative World Model for Autonomous Driving

2023-09-29 · Anthony Hu, Lloyd Russell, Hudson Yeo et al. · arXiv:2309.17080

GAIA-1 is a generative world model for driving that takes video, text and action inputs, casts world modelling as next-token prediction over discrete tokens, and decodes with a video diffusion model to produce realistic driving scenes with control over ego-vehicle behaviour and scene features.

DrivingEditor pick

DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

2023-09-18 · Xiaofeng Wang, Zheng Zhu, Guan Huang et al. · arXiv:2309.09777

DriveDreamer is a diffusion-based world model learned from real-world driving data (evaluated on nuScenes) that generates controllable driving videos and future driving policies, trained in two stages (structured traffic constraints first, then future-state prediction).

Drivingauto-summary

Mastering Diverse Domains through World Models

2023-01-10 · Danijar Hafner, Jurgis Pasukonis, Jimmy Ba et al. · arXiv:2301.04104

DreamerV3 is a general model-based RL algorithm that learns a world model and improves behavior by imagining future scenarios; with a single configuration it outperforms specialized methods across over 150 tasks, and is the first to collect diamonds in Minecraft from scratch without human data.

Latent Dynamicscodeauto-summary

Transformers are Sample-Efficient World Models

2022-09-01 · Vincent Micheli, Eloi Alonso, François Fleuret · arXiv:2209.00588

IRIS is a data-efficient RL agent that learns inside a world model built from a discrete autoencoder plus an autoregressive Transformer, and reaches a mean human-normalized score of 1.046 on Atari 100k (about two hours of gameplay).

Latent Dynamicscodeauto-summary

DayDreamer: World Models for Physical Robot Learning

2022-06-28 · Philipp Wu, Alejandro Escontrela, Danijar Hafner et al. · arXiv:2206.14176

DayDreamer applies the Dreamer algorithm to four physical robots that learn online in the real world without simulators, using the same hyperparameters across tasks: a quadruped, two robot arms and a wheeled robot.

Roboticsauto-summary