RECON improves model-based imitation learning by separating conservative policy learning from active real-environment data collection. It trains a main (conservative) policy for task execution and an explorer optimized using epistemic uncertainty conditioned on recoverability estimated from multi-step imagination with the main policy, prioritizing uncertain yet recoverable dynamics. Experiments across locomotion, navigation, and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness.
MultiEcho treats frozen world models as experimental systems by estimating response laws from controlled counterfactual interventions. Using “three-reference” forward fits and reverse readouts, the authors predict complete intervention responses and recover intervention parameters across nine simulated physical systems and seven frozen model configurations, then test applicability and physical correspondence via temporal/visual/material intervention variants.
KoopCell is a Koopman-based generative framework for learning continuous single-cell population dynamics from temporally sparse, unpaired distribution snapshots. It jointly learns latent coordinates and linear latent dynamics guided by Koopman–Mori–Zwanzig theory, with convergence guarantees from a weak continuity equation. KoopCell-M adds non-Markovian memory via Markovian embedding to model branching in latent dynamics.
2026-09-30 · Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi et al. · arXiv:2609.33487
The authors run a controlled sweep over spatio-temporal EEG masking geometry for 58 masked-prediction foundation models trained with MAE and JEPA, then evaluate all models on 12 OpenEEGBench datasets using linear probing. Both frameworks agree on an optimal mask configuration and shared failure modes; performance is robust outside the best region, while JEPA has an additional failure mode (“bias-inflation collapse”).
2026-09-30 · Tamim Zoabi, Ameen Ali, Lior Wolf · arXiv:2609.33497
H-JEPA is an action-conditioned, reconstruction-free world model that separates a wide perceptual code from a fixed orthonormal control state slice. The control state inherits the code’s covariance, then evolves with phase-conditioned dissipative port-Hamiltonian dynamics. Port-inverse consistency (PIC) ties the action readout to the input port transpose, shown to equal a rollout error projection reweighting prediction error. The authors report matching or exceeding reconstruction-free baselines on four pixel-based control benchmarks in ≤10 epochs.
2026-09-30 · Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt · arXiv:2609.33563
MA-JEPA introduces a stochastic joint-embedding predictive world model for multi-agent RL with centralized training and decentralized execution. It replaces observation reconstruction with prediction of target observation embeddings, using a categorical latent state and a causal Transformer. A training-only joint predictor conditions on all agents’ local states/actions to predict each agent’s next embedding, which is passed through the same local posterior as in execution for actor-critic learning from latent imagination.
The authors introduce Observer State-Space Models (OSSMs) for time-series forecasting. OSSMs treat the observed input sequence as measurements correcting an internal latent state via an observer, while the latent transition is autonomous and shared across context and forecasting intervals. They claim OSSM improves forecasting versus corresponding SSMs under the same parameter count and training setup.
2026-09-30 · Boyuan Zhang, Yingjun Du, Xiantong Zhen et al. · arXiv:2609.33595
The authors introduce SALT, an action-conditioned joint-embedding latent dynamics model with state-affine transitions and recursive rollout training. They argue one-step prediction error can fail to predict planning quality because planning composes transitions recursively, transforming introduced errors. Across four visual planning environments, SALT has higher one-step prediction error than LeWM but higher closed-loop success.
2026-09-30 · Teng Cao, Yu Deng, Quentin Delfosse et al. · arXiv:2609.33728
ALDER discovers and revises explicit, equation-based world models by acting in the environment. It proposes parametric equation structures, fits coefficients, and uses an independent verifier on held-out data. When multiple hypotheses fit, an experiment selector queries cost- and safety-aware interventions where predictions disagree. Counterexamples update evidence and drive structural revisions; validated equations are used for prediction and control via an inverse problem.
2026-09-30 · Yuhao Li, Louie Hong Yao, Tianyi Shi et al. · arXiv:2609.33804
MinkowskiPE proposes Minkowski positional encoding for spatiotemporal perception. It assigns each token a spacetime coordinate and uses the relative spacetime displacement between tokens to parameterize Lorentz transformations applied to query/key features, making attention depend on position only through relative displacement and be translation-invariant. It retains the standard dot-product attention interface and is compatible with efficient attention implementations.
2026-09-30 · Yuheng Qiao, Ziran Wei, Xiaohan Wang et al. · arXiv:2609.33832
The authors treat a world-action model’s visual predictions as a goal-conditioned visual proposal and use a frozen action-conditioned world model to predict action consequences. They compute dense feedback from cross-model prediction consistency plus terminal goal alignment, then use Flow Policy Optimization to update only the action head. Across four real-world UR5 tasks, they report mean success increasing from 43.4% to 75.1%, without online robot interaction or task-specific reward models.
2026-09-30 · Jifan Li, Ning Ning · arXiv:2609.33844
ViBR-WM is a visual world model that forecasts in a DINOv2-based latent representation and uses Bayesian regression to predict joint visual–physical states recursively and physical targets directly. It combines trend, seasonal, and cycle modules with interpretable regression and Bayesian variable selection, averaging over predictor subsets via posterior probabilities and modeling parameter uncertainty and future disturbances.
QwenGyre is an end-to-end framework for x Long-horizon online RL with black-box LLM agent harnesses. It uses an elastic scheduler to elastically reallocate GPUs between rollout and training without interrupting live executions, and a trajectory processor that reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training cost. On NL2RepoBench with Qwen3.8 2.4T, it improves score in 48 steps and reports end-to-end speedups over Colocate and Async.
2026-09-30 · Thomas Walker, Randall Balestriero, Richard Baraniuk · arXiv:2609.33940
The authors propose behavioral monitoring signals for JEPA world models using centroids defined as sub-component Jacobian row-sums. Centroids can be computed via Jacobian vector products, characterize internal input-space geometry, and yield saliency maps. On Push-T and TwoRoom under distribution shift, a structural dissociation (encoder goal represented, predictor unresponsive) predicts planning failure before actions; centroid-based methods outperform activation- and reconstruction-based shift detectors, enabling pre-execution goal resampling.
2026-09-30 · Alexander Detkov, Matt Thomson · arXiv:2609.34058
The authors study whether models learn global constraints and propagate their consequences from local transition data. In controlled monoid worlds, next-state training fits paths but fails to propagate non-trivial constraints, while compositional training (hiding intermediates) achieves 96% average accuracy on inverse/commutativity/composition. Generalization declines sharply with proof depth; longer compositional path length improves deeper propagation and improves related results in vision and language.
AD-E2E-JEPA is an action-conditioned JEPA world model for end-to-end autonomous driving that enables goal-conditioned zero-shot planning without training a driving policy. It adds a SIGReg-regularized learnable projector that compresses patch embeddings, reducing planning latency (0.8s for an 8-frame rollout over 256 trajectories) while retaining planning performance on NAVSIMv2.
WorldGraph studies graph world modeling (GWM) where the evolving graph itself is the world dynamics. It formulates latent graph states and heterogeneous transition prediction over node-, edge-, and graph-level changes, builds the GWM-Zero benchmark (8 temporal graph datasets), and proposes WorldGraph with a state-aware graph transformer and a transition-aware GRPO (dynamic grouping, structure-aware verifiable rewards).
2026-09-30 · Ziyao Zeng, Xiatao Sun, Hao Wang et al. · arXiv:2609.34286
DTWM is a video world model for future-frame prediction of egocentric human manipulation using both RGB video and dexterous tactile glove signals from each hand. It conditions a pretrained video diffusion transformer by injecting a zero-initialized tactile residual at each hand’s location in the video tokens with block-causal masking so predicted frames cannot use future info. Compared to a matched vision-only baseline, it reduces hand-motion underestimation (23% to 9%) and hand-region perceptual error.
2026-09-30 · Luzhe Huang, Lei Chu, Jingyi Liang et al. · arXiv:2609.34375
LRC-JEPA is a compact JEPA world model that routes information into two streams: a predictive latent z for action-conditioned dynamics and planning, and a residual context embedding u for reconstruction. Only z is rolled out by the dynamics model at test time. The authors report improved planning success in simulated control and better Bridge-v2 offline action-recovery while using a 5.5M active-parameter encoder and faster planning.
2026-09-30 · Haojie Yang, Ran Su · arXiv:2609.34391
P2P predicts unpaired single-cell perturbation response at the population level by learning shared condition effects from two stochastic “views” of the same control/perturbed condition. It encodes control sets with a permutation-invariant set encoder, perturbations with a structured encoder, and blends an empirical memory with a neural residual to output population mean and gene-wise variance. Trained with cross-view supervision (no cell correspondences).