Latent Dynamics · 2026-09-30
ViBR-WM: Visual Bayesian Regression for World Modeling
Jifan Li, Ning Ning
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
ViBR-WM is a visual world model that forecasts in a DINOv2-based latent representation and uses Bayesian regression to predict joint visual–physical states recursively and physical targets directly. It combines trend, seasonal, and cycle modules with interpretable regression and Bayesian variable selection, averaging over predictor subsets via posterior probabilities and modeling parameter uncertainty and future disturbances.
Why it matters
The authors target forecasting with uncertainty in world-model settings that use both images and numerical signals. ViBR-WM builds a structured Bayesian architecture over a fixed visual state space, using Bayesian variable selection and posterior prediction to quantify uncertainty while improving physical-target forecasting error versus Temporal Straightening, ConvLSTM, PredRNN, and SimVP across four tasks.
Method
- Encode RGB frames with a frozen DINOv2 ViT-S/14 backbone, project TS-trained features with learned projector, and compress via two-stage PCA into visual coordinates.
- Model state changes with Bayesian regression of changes in a joint visual–physical state, optionally augmented by trend, seasonal, and cycle modules with cycle damping.
- Generate forecasts by posterior prediction: sample regression parameters and disturbances; for direct targets, predict each horizon without feeding back prior predictions.
Abstract (from arXiv)
Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle dynamics. Visual compression reduces representation dimension, while Bayesian variable selection reduces active regression dimension. Posterior prediction combines forecasts across predictor subsets using their posterior probabilities as weights and accounts for parameter uncertainty and future disturbances. The model forecasts joint visual--physical states recursively and physical targets directly. Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task. Repeated fitting and resampling support these overall gains.
Related papers
- MeteoVerse: Unified Weather-Controllable Video World Model
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
- Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
- Adaptive Latent Capacity for World Models
- Abductive World Modeling via Causal Representation Learning