Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Do World Models Learn Global Understanding?

Alexander Detkov, Matt Thomson

arXiv:2609.34058PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors study whether models learn global constraints and propagate their consequences from local transition data. In controlled monoid worlds, next-state training fits paths but fails to propagate non-trivial constraints, while compositional training (hiding intermediates) achieves 96% average accuracy on inverse/commutativity/composition. Generalization declines sharply with proof depth; longer compositional path length improves deeper propagation and improves related results in vision and language.

Why it matters

The paper formalizes “global understanding” as learning constraints and propagating implied held-out transitions. It shows that training objectives change whether models integrate local information into coherent global structure, across attention, recurrent, and state-space models, and also in visual world models and Wikidata-finetuned LLMs. The proof-depth analysis provides a way to quantify how far consequences are propagated when further inferences are required.

Method

  • Construct monoid worlds: states with action-defined transitions, plus specified equality constraints over action sequences; create entailed hold-out transitions from observed transitions + global constraint.
  • Train on length-(T) paths of observed transitions under two objectives: next-state (intermediate states shown) vs compositional (hide intermediates; predict composition internally).
  • Measure constraint propagation via entailed hold-outs and quantify inference complexity using proof depth: minimum inference rounds to derive a held-out transition; test generalization vs proof depth and composition length T.

Limitation

Periodicity constraints are not propagated by next-state or compositional training (at T=2 or 4 in monoid worlds), and all transformers trained compositionally on periodicity worlds with A^k=1 for k=3,4,5 and composition lengths 1<=T<=8 fail to propagate the periodicity constraint.

Abstract (from arXiv)

AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.

Related papers