Video World Models · 2026-09-30
Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
The authors evaluate physical prediction in video world models at “event boundaries,” where release/impact is shown but its consequence is withheld. Using 62 controlled real free-fall recordings (124 event-anchored clips) and human trajectory prediction, they find Runway and Veo often produce release and impact events above 93% but with late onset, while Cosmos-Predict-2.5 and MAGI-1 frequently suppress measurable consequences. They also find that plausible timing does not guarantee physically consistent motion.
Why it matters
Event-anchored evaluation directly tests whether a model initiates and carries through the consequence of an observed physical event, rather than inferring anticipation indirectly from reference similarity or physics scores. The authors separate consequence production, temporal placement, and physical realization, showing distinct failure modes and highlighting that one physical metric (e.g., timing) does not imply overall physical consistency.
Method
- Introduce an event-anchored, outcome-blind protocol that withholds everything after an annotated release or impact, separately measuring consequence production, temporal placement, and physical realization.
- Create a controlled real-world free-fall dataset with 62 recordings split into 124 release- and impact-anchored clips, including release/impact/rest annotations and metric trajectories.
- Run a human study (15 participants, 20 conditions) predicting the expected consequence and drawing trajectories from a single event-anchored frame, then compare human and model support and geometry.
Limitation
The authors note that comparisons with Runway and Veo should be interpreted as alignment with a minimal-evidence human reference rather than as a controlled test of context length, since video-conditioned models receive additional temporal context while Runway/Veo use only the anchor frame.
Abstract (from arXiv)
Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.
Related papers
- MultiEcho: An Experimental Science of Learned Worlds
- Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
- World2Motion: Turning Video World Models into 3D Human Motion Generators
- One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions
- ViBR-WM: Visual Bayesian Regression for World Modeling