Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Abductive World Modeling via Causal Representation Learning

Ziqi Liu, Songhan Yang, Linfan Zhou, Jiatong Liu, Lijun Peng, Long Wan, Yinqi Bai

arXiv:2609.36985PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors propose Abductive World Modeling (AWM), which learns structured causal representations by “predict forward, then abduce backward.” Using the Hierarchical Abductive State Pyramid (HASP), AWM infers a hierarchical latent state with Entity, Dynamic, and Relation components by jointly reasoning over the current observation and its predicted future. Experiments on physical prediction, causal reasoning, and action understanding show consistent gains over a V-JEPA 2 backbone-only baseline.

Why it matters

The paper targets a limitation of latent-space world models that predict future states without explicitly capturing latent causes for why/how the world changes. By treating the predicted future as evidence and inferring structured latent factors (what exists, how it changes, how entities interact), AWM aims to provide a more usable interface for downstream reasoning than an entangled predictive representation.

Method

  • AWM predicts a future latent state from the observed context, then uses this predicted future as evidence to infer abductive latent causes.
  • HASP constructs a hierarchical abductive state with three components: Entity (what exists), Dynamic (how entities change over time), and Relation (how entities interact).
  • Training freezes the predictive encoder and latent predictor; only HASP and task-specific readouts are optimized, aligning supervision to each state’s native granularity.
Abstract (from arXiv)

The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.

Related papers