Video World Models · 2026-09-30
MeteoVerse: Unified Weather-Controllable Video World Model
Renlong Wu, Guanqiao Wang, Xuan Shang, Yin Hanming, Xiaoxiao Sheng, Tianyu Huang, Hui Li, Wangmeng Zuo
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
MeteoVerse is a unified weather-controllable video world model for camera-controlled future prediction from either sunny or adverse-weather inputs. It explicitly estimates observed and target weather states (rain/snow/fog), derives a weather transition, and uses a transition-aware mixture of weather experts (MeteoMoE) to inject residual weather features into a frozen video backbone. It also introduces a dataset of over 50K real-world weather clips with sunny pseudo-pairs and weather-intensity annotations.
Why it matters
Prior controllable video models often leave weather transition implicit, forcing the generation backbone to infer weather evolution alongside scene and camera dynamics, which reduces precision in weather control. By modeling the transition between observed and desired weather states and translating it via MeteoMoE into residual, category-specific weather features, MeteoVerse aims to improve weather controllability while maintaining competitive scene consistency and camera-control performance.
Method
- Given a single image (sunny or adverse), a weather-free scene description, a target-weather instruction, and a camera trajectory, a weather-state predictor estimates continuous observed and target intensities over rain/snow/fog and computes a transition vector s_trans = s_tar − s_src.
- A transition-aware mixture of weather experts (MeteoMoE) uses category-specific rain/snow/fog experts, modulates expert responses by the transition components, fuses them, and residually injects transition-aware features into a DiT video backbone.
- Training uses a pretrained LingBot-World backbone with MeteoMoE and LoRA adapters optimized with a flow-matching objective plus frame-wise reconstruction and perceptual losses. The predictor is LoRA-fine-tuned and frozen at inference.
Abstract (from arXiv)
Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determined not only by changes in viewpoint and object dynamics, but also by environmental conditions such as weather, which can substantially alter scene appearance and visibility. Modeling such realistic weather evolution is challenging because the required weather modification depends jointly on the observed and desired weather states. Depending on their relation, the model may need to preserve, introduce, or remove a weather effect. Existing video world models typically leave this weather transition implicit, forcing the generation backbone to infer weather evolution together with scene dynamics and camera motion, which leads to imprecise weather control. To address this limitation, we propose MeteoVerse, a unified weather-controllable video world model that generates future videos from a single sunny or adverse-weather image, conditioned on a weather-free scene description, a target-weather instruction, and a camera trajectory. Rather than conditioning only on the desired weather, MeteoVerse explicitly estimates the observed and target weather states and represents the required weather transition. A transition-aware mixture of weather experts then translates this transition into category-specific residual weather features, unifying weather preservation, introduction, and removal while enabling fine-grained control over introduced weather intensity. We further construct the MeteoVerse dataset with over 50K real-world weather video clips, generated sunny counterparts, disentangled scene and weather descriptions, weather-intensity annotations, and camera trajectories. Extensive experiments demonstrate substantially improved weather controllability while retaining competitive scene consistency and camera-control performance.
Related papers
- ViBR-WM: Visual Bayesian Regression for World Modeling
- Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
- World2Motion: Turning Video World Models into 3D Human Motion Generators
- Honeycomb: Constant-Size Scene Memory Representation for Video World Models
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability