Pub-AI: AI in Science

Latent Dynamics · 2026-09-30

Do-JEPA: From Masking to Intervention in Latent World Models

Hossein Resani, Javen Qinfeng Shi

arXiv:2609.37378PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

Do-JEPA trains latent world models using paired simulator rollouts from the same saved state under an action a and a reference action a∅, supervising the latent difference Δz. The objective decomposes into effect, support, propagation, and invariance losses. In pixels and learned slots, the authors report improved action-causal (responsive) effect prediction versus masking-style baselines, and in fine-tuning reduces the factual-cost of training from scratch.

Why it matters

The paper argues that standard latent prediction objectives do not separate what an action caused from what co-occurred. By changing what the data represents—running counterfactual action branches from an exactly restorable simulator state—the method provides direct training signals for where an action enters, how its effect propagates, and what must remain invariant under interventions and context shifts.

Method

  • From one saved simulator state, run dynamics twice for k steps under action a and reference action a∅, then train the model to predict the paired latent effect difference Δz=z^a−z^{a∅}.
  • Use Do-JEPA losses: effect loss; support loss (where the action enters); propagation loss (how effect travels via onset-based edge labels); invariance losses (non-responders and context twins).
  • At test time, given observation history and an action, run the predictor for a and a∅ to obtain the effect prediction Δẑ; planning uses only the factual prediction.

Limitation

The method’s losses “need an exactly restorable simulator (all losses), object-aligned slots (L_sup, L_edge, L_inv), and a way to change the context without changing the physics (L_ctx).”

Abstract (from arXiv)

Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $\Delta z=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.

Related papers