Pub-AI: AI in Science

Video World Models

Generative video models used as world simulators. (16 papers)

Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

2026-09-30 · Jiaming Zhang, Wu Yang, Shuai Tao et al. · arXiv:2609.31938

Cache-aware Conv3D lowering for embedded generative video VAEs: supported causal Conv3D calls are expressed as batched spatial Conv2D while preserving pretrained weights, causal-cache semantics, convolution parameters, bias placement, and output layout. On 64-GB Jetson AGX Orin in the Cosmos3-Edge image-to-video pipeline, the authors report ~7.32× VAE-decoder speedup and 2.21× end-to-end speedup at 25 frames.

Video World Modelsauto-summary

Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State

2026-09-30 · Tamim Zoabi, Ameen Ali, Lior Wolf · arXiv:2609.33497

H-JEPA is an action-conditioned, reconstruction-free world model that separates a wide perceptual code from a fixed orthonormal control state slice. The control state inherits the code’s covariance, then evolves with phase-conditioned dissipative port-Hamiltonian dynamics. Port-inverse consistency (PIC) ties the action readout to the input port transpose, shown to equal a rollout error projection reweighting prediction error. The authors report matching or exceeding reconstruction-free baselines on four pixel-based control benchmarks in ≤10 epochs.

Video World Modelsauto-summary

MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception

2026-09-30 · Yuhao Li, Louie Hong Yao, Tianyi Shi et al. · arXiv:2609.33804

MinkowskiPE proposes Minkowski positional encoding for spatiotemporal perception. It assigns each token a spacetime coordinate and uses the relative spacetime displacement between tokens to parameterize Lorentz transformations applied to query/key features, making attention depend on position only through relative displacement and be translation-invariant. It retains the standard dot-product attention interface and is compatible with efficient attention implementations.

Video World Modelsauto-summary

Dexterous Tactile World Model

2026-09-30 · Ziyao Zeng, Xiatao Sun, Hao Wang et al. · arXiv:2609.34286

DTWM is a video world model for future-frame prediction of egocentric human manipulation using both RGB video and dexterous tactile glove signals from each hand. It conditions a pretrained video diffusion transformer by injecting a zero-initialized tactile residual at each hand’s location in the video tokens with block-causal masking so predicted frames cannot use future info. Compared to a matched vision-only baseline, it reduces hand-motion underestimation (23% to 9%) and hand-region perceptual error.

Video World Modelsauto-summary

Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models

2026-09-30 · Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen et al. · arXiv:2609.34677

The authors propose Future-Aware Recall (FAR), a framework for episodic memory access in world models. FAR learns which past memories to retrieve and which retrieval cues (time, pose, vision, audio) to trust by supervising a future-blind retriever using future-aware predictive utility during training (approximated by negative diffusion prediction loss).

Video World Modelsauto-summary

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

2026-09-30 · Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico) et al. · arXiv:2609.36531

The authors evaluate physical prediction in video world models at “event boundaries,” where release/impact is shown but its consequence is withheld. Using 62 controlled real free-fall recordings (124 event-anchored clips) and human trajectory prediction, they find Runway and Veo often produce release and impact events above 93% but with late onset, while Cosmos-Predict-2.5 and MAGI-1 frequently suppress measurable consequences. They also find that plausible timing does not guarantee physically consistent motion.

Video World Modelsauto-summary

MeteoVerse: Unified Weather-Controllable Video World Model

2026-09-30 · Renlong Wu, Guanqiao Wang, Xuan Shang et al. · arXiv:2609.36810

MeteoVerse is a unified weather-controllable video world model for camera-controlled future prediction from either sunny or adverse-weather inputs. It explicitly estimates observed and target weather states (rain/snow/fog), derives a weather transition, and uses a transition-aware mixture of weather experts (MeteoMoE) to inject residual weather features into a frozen video backbone. It also introduces a dataset of over 50K real-world weather clips with sunny pseudo-pairs and weather-intensity annotations.

Video World Modelsauto-summary

World2Motion: Turning Video World Models into 3D Human Motion Generators

2026-09-30 · Tu Fangyuan, Xiangyue Zhang, Yiyi Cai et al. · arXiv:2609.37004

World2Motion adapts Cosmos 3 into a single-stage generator that produces scene-aware 3D human motion and corresponding video from a single image and a text prompt. The authors build a hybrid training dataset (synthetic video–motion pairs plus real videos paired with estimated 3D motion) and introduce a shift-decoupled noise schedule to reduce motion jitter. On a multi-source interaction benchmark, it improves motion–text alignment and scene interaction and matches two-stage interaction success with ~3.3× faster inference.

Video World Modelsauto-summary

Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

2026-09-30 · Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn et al. · arXiv:2609.37107

Waypoint-1.5 is a real-time diffusion video world model for interactive generation on consumer GPUs. It uses a single-stream causal Diffusion Transformer over TAEHV 1.5 latents and is conditioned on full keyboard+mouse inputs. Trained on 100,000 hours of control-aligned gameplay, it produces playable rollouts with a KV-cache runtime and distillation that reduces denoising to 4 steps per frame. It is benchmarked via latent FPS throughput.

Video World Modelsauto-summary

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

2026-09-30 · Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen et al. · arXiv:2609.37690

Honeycomb is a video world model that keeps persistent scene memory in a fixed-size HexMemory. HexMemory stores scene features as six fixed-dimension low-rank planes (three spatial, three spatiotemporal). A feed-forward writer updates only new video chunks by warping prior planes to expanded bounds and fusing with confidence-weighted pooling plus a learned residual. A reader reconstructs latent features from HexMemory to condition generation; the method avoids per-scene optimization and full-history reprocessing.

Video World Modelsauto-summary

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

2026-09-30 · Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as et al. · arXiv:2609.37708

BRAID (Bilevel Representations for Agent Interaction Dynamics) is a hierarchical sequential latent-variable model for generative multi-person social motion. It uses a group latent for shared interaction dynamics and person latents for individual behavior conditioned on the evolving group context, enabling coherent generation and compact social-state vectors for full, sparse, or partial observations.

Video World Modelsauto-summary

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

2026-09-30 · Yu Xu, Yuxin Zhang, Xiao Yang et al. · arXiv:2609.38140

The authors argue that token-wise load-balanced MoE routing forms a “uniformity trap” for video, scattering coherent spatiotemporal tokens across experts and causing fragmentation and distortion. They propose SplitMoE, splitting experts into semantic and generic roles with prototype-guided semantic routing and pull-push regularization, reporting faster convergence and improved routing coherence and video generation quality under matched activated-parameter budgets.

Video World Modelsauto-summary

LongLive-Plug: Once-for-All Distillation for Video Generation

2026-09-30 · Shuai Yang, Luozhou Wang, Wei Huang et al. · arXiv:2609.38154

LongLive-Plug introduces a once-for-all distillation framework that learns reusable LoRA adapters on a base video diffusion model. The adapters cover single-pass classifier-free guidance, few-step (four-step) sampling, and long-context error correction for causal autoregressive generation, enabling training-free plug-and-play deployment across compatible downstream models.

Video World Modelsauto-summary

Cosmos World Foundation Model Platform for Physical AI

2025-01-07 · NVIDIA, :, Niket Agarwal et al. · arXiv:2501.03575

Cosmos is NVIDIA's platform for 'world foundation models' for Physical AI: a video curation pipeline, pre-trained diffusion and autoregressive world models, examples of post-training for downstream tasks, and video tokenizers, released open-source with open-weight models.

Video World ModelscodeEditor pick

Genie: Generative Interactive Environments

2024-02-23 · Jake Bruce, Michael Dennis, Ashley Edwards et al. · arXiv:2402.15391

Genie is an 11B-parameter generative interactive environment trained without action labels on unlabelled internet videos; a spatiotemporal video tokenizer, autoregressive dynamics model and latent action model let users act frame by frame in generated worlds.

Video World Modelsauto-summary

Learning Interactive Real-World Simulators

2023-10-09 · Sherry Yang, Yilun Du, Kamyar Ghasemipour et al. · arXiv:2310.06114

UniSim learns a universal simulator of real-world interaction by orchestrating diverse datasets (image, robotics, navigation) so it can simulate the visual outcome of both high-level instructions and low-level controls; policies trained purely inside it deploy zero-shot on real robots.

Video World Modelsauto-summary