Pub-AI: AI in Science

Digest

Reset

126 papers

Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning

2026-09-30 · Xuanlin Chen, Ziyue Wang, Xunlan Zhou et al. · arXiv:2609.33336

RECON improves model-based imitation learning by separating conservative policy learning from active real-environment data collection. It trains a main (conservative) policy for task execution and an explorer optimized using epistemic uncertainty conditioned on recoverability estimated from multi-step imagination with the main policy, prioritizing uncertain yet recoverable dynamics. Experiments across locomotion, navigation, and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness.

Drivingauto-summary

MultiEcho: An Experimental Science of Learned Worlds

2026-09-30 · Meng Zhu, Airui Zhang · arXiv:2609.33347

MultiEcho treats frozen world models as experimental systems by estimating response laws from controlled counterfactual interventions. Using “three-reference” forward fits and reverse readouts, the authors predict complete intervention responses and recover intervention parameters across nine simulated physical systems and seven frozen model configurations, then test applicability and physical correspondence via temporal/visual/material intervention variants.

Latent Dynamicsauto-summary

KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots

2026-09-30 · Wanfeng Lu, Yutong Zhang, Keyi Zhou et al. · arXiv:2609.33350

KoopCell is a Koopman-based generative framework for learning continuous single-cell population dynamics from temporally sparse, unpaired distribution snapshots. It jointly learns latent coordinates and linear latent dynamics guided by Koopman–Mori–Zwanzig theory, with convergence guarantees from a weak continuity equation. KoopCell-M adds non-Markovian memory via Markovian embedding to model branching in latent dynamics.

JEPAauto-summary

What masking geometry works best for EEG foundation models?

2026-09-30 · Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi et al. · arXiv:2609.33487

The authors run a controlled sweep over spatio-temporal EEG masking geometry for 58 masked-prediction foundation models trained with MAE and JEPA, then evaluate all models on 12 OpenEEGBench datasets using linear probing. Both frameworks agree on an optimal mask configuration and shared failure modes; performance is robust outside the best region, while JEPA has an additional failure mode (“bias-inflation collapse”).

JEPAauto-summary

Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State

2026-09-30 · Tamim Zoabi, Ameen Ali, Lior Wolf · arXiv:2609.33497

H-JEPA is an action-conditioned, reconstruction-free world model that separates a wide perceptual code from a fixed orthonormal control state slice. The control state inherits the code’s covariance, then evolves with phase-conditioned dissipative port-Hamiltonian dynamics. Port-inverse consistency (PIC) ties the action readout to the input port transpose, shown to equal a rollout error projection reweighting prediction error. The authors report matching or exceeding reconstruction-free baselines on four pixel-based control benchmarks in ≤10 epochs.

Video World Modelsauto-summary

MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

2026-09-30 · Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt · arXiv:2609.33563

MA-JEPA introduces a stochastic joint-embedding predictive world model for multi-agent RL with centralized training and decentralized execution. It replaces observation reconstruction with prediction of target observation embeddings, using a categorical latent state and a causal Transformer. A training-only joint predictor conditions on all agents’ local states/actions to predict each agent’s next embedding, which is passed through the same local posterior as in execution for actor-critic learning from latent imagination.

JEPAauto-summary

Correct then Forecast: Observer State-Space Models for Time Series Forecasting

2026-09-30 · Alexis-Raja Brachet, Guillaume Clavier--Fr\'emond, Abdelhakim Ziani et al. · arXiv:2609.33566

The authors introduce Observer State-Space Models (OSSMs) for time-series forecasting. OSSMs treat the observed input sequence as measurements correcting an internal latent state via an observer, while the latent transition is autonomous and shared across context and forecasting intervals. They claim OSSM improves forecasting versus corresponding SSMs under the same parameter count and training setup.

Latent Dynamicsauto-summary

Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning

2026-09-30 · Boyuan Zhang, Yingjun Du, Xiantong Zhen et al. · arXiv:2609.33595

The authors introduce SALT, an action-conditioned joint-embedding latent dynamics model with state-affine transitions and recursive rollout training. They argue one-step prediction error can fail to predict planning quality because planning composes transitions recursively, transforming introduced errors. Across four visual planning environments, SALT has higher one-step prediction error than LeWM but higher closed-loop success.

Latent Dynamicscodeauto-summary

ALDER: Discovering the Laws of a World by Acting in It

2026-09-30 · Teng Cao, Yu Deng, Quentin Delfosse et al. · arXiv:2609.33728

ALDER discovers and revises explicit, equation-based world models by acting in the environment. It proposes parametric equation structures, fits coefficients, and uses an independent verifier on held-out data. When multiple hypotheses fit, an experiment selector queries cost- and safety-aware interventions where predictions disagree. Counterexamples update evidence and drive structural revisions; validated equations are used for prediction and control via an inverse problem.

Drivingauto-summary

MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception

2026-09-30 · Yuhao Li, Louie Hong Yao, Tianyi Shi et al. · arXiv:2609.33804

MinkowskiPE proposes Minkowski positional encoding for spatiotemporal perception. It assigns each token a spacetime coordinate and uses the relative spacetime displacement between tokens to parameterize Lorentz transformations applied to query/key features, making attention depend on position only through relative displacement and be translation-invariant. It retains the standard dot-product attention interface and is compatible with efficient attention implementations.

Video World Modelsauto-summary

Achieve What You Imagined: Learning to Align Actions with Visual Plans

2026-09-30 · Yuheng Qiao, Ziran Wei, Xiaohan Wang et al. · arXiv:2609.33832

The authors treat a world-action model’s visual predictions as a goal-conditioned visual proposal and use a frozen action-conditioned world model to predict action consequences. They compute dense feedback from cross-model prediction consistency plus terminal goal alignment, then use Flow Policy Optimization to update only the action head. Across four real-world UR5 tasks, they report mean success increasing from 43.4% to 75.1%, without online robot interaction or task-specific reward models.

Drivingauto-summary

ViBR-WM: Visual Bayesian Regression for World Modeling

2026-09-30 · Jifan Li, Ning Ning · arXiv:2609.33844

ViBR-WM is a visual world model that forecasts in a DINOv2-based latent representation and uses Bayesian regression to predict joint visual–physical states recursively and physical targets directly. It combines trend, seasonal, and cycle modules with interpretable regression and Bayesian variable selection, averaging over predictor subsets via posterior probabilities and modeling parameter uncertainty and future disturbances.

Latent Dynamicsauto-summary

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

2026-09-30 · Weiqi Wang, Yuxin Zhou, Mouxiang Chen et al. · arXiv:2609.33848

QwenGyre is an end-to-end framework for x Long-horizon online RL with black-box LLM agent harnesses. It uses an elastic scheduler to elastically reallocate GPUs between rollout and training without interrupting live executions, and a trajectory processor that reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training cost. On NL2RepoBench with Qwen3.8 2.4T, it improves score in 48 steps and reports end-to-end speedups over Colocate and Async.

Drivingauto-summary

Behavioral Monitoring of JEPA World Models with Jacobian Centroids

2026-09-30 · Thomas Walker, Randall Balestriero, Richard Baraniuk · arXiv:2609.33940

The authors propose behavioral monitoring signals for JEPA world models using centroids defined as sub-component Jacobian row-sums. Centroids can be computed via Jacobian vector products, characterize internal input-space geometry, and yield saliency maps. On Push-T and TwoRoom under distribution shift, a structural dissociation (encoder goal represented, predictor unresponsive) predicts planning failure before actions; centroid-based methods outperform activation- and reconstruction-based shift detectors, enabling pre-execution goal resampling.

JEPAauto-summary

Do World Models Learn Global Understanding?

2026-09-30 · Alexander Detkov, Matt Thomson · arXiv:2609.34058

The authors study whether models learn global constraints and propagate their consequences from local transition data. In controlled monoid worlds, next-state training fits paths but fails to propagate non-trivial constraints, while compositional training (hiding intermediates) achieves 96% average accuracy on inverse/commutativity/composition. Generalization declines sharply with proof depth; longer compositional path length improves deeper propagation and improves related results in vision and language.

Latent Dynamicsauto-summary

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

2026-09-30 · Haoran Zhu, Wancong Zhang, Yann LeCun et al. · arXiv:2609.34085

AD-E2E-JEPA is an action-conditioned JEPA world model for end-to-end autonomous driving that enables goal-conditioned zero-shot planning without training a driving policy. It adds a SIGReg-regularized learnable projector that compresses patch embeddings, reducing planning latency (0.8s for an 8-frame rollout over 256 trajectories) while retaining planning performance on NAVSIMv2.

Drivingcodeauto-summary

WorldGraph: Graph-Native World Modeling

2026-09-30 · Zezhong Ding, Yipeng Li, Xike Xie · arXiv:2609.34159

WorldGraph studies graph world modeling (GWM) where the evolving graph itself is the world dynamics. It formulates latent graph states and heterogeneous transition prediction over node-, edge-, and graph-level changes, builds the GWM-Zero benchmark (8 temporal graph datasets), and proposes WorldGraph with a state-aware graph transformer and a transition-aware GRPO (dynamic grouping, structure-aware verifiable rewards).

JEPAauto-summary

Dexterous Tactile World Model

2026-09-30 · Ziyao Zeng, Xiatao Sun, Hao Wang et al. · arXiv:2609.34286

DTWM is a video world model for future-frame prediction of egocentric human manipulation using both RGB video and dexterous tactile glove signals from each hand. It conditions a pretrained video diffusion transformer by injecting a zero-initialized tactile residual at each hand’s location in the video tokens with block-causal masking so predicted frames cannot use future info. Compared to a matched vision-only baseline, it reduces hand-motion underestimation (23% to 9%) and hand-region perceptual error.

Video World Modelsauto-summary

LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models

2026-09-30 · Luzhe Huang, Lei Chu, Jingyi Liang et al. · arXiv:2609.34375

LRC-JEPA is a compact JEPA world model that routes information into two streams: a predictive latent z for action-conditioned dynamics and planning, and a residual context embedding u for reconstruction. Only z is rolled out by the dynamics model at test time. The authors report improved planning success in simulated control and better Bridge-v2 offline action-recovery while using a 5.5M active-parameter encoder and faster planning.

JEPAauto-summary

P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction

2026-09-30 · Haojie Yang, Ran Su · arXiv:2609.34391

P2P predicts unpaired single-cell perturbation response at the population level by learning shared condition effects from two stochastic “views” of the same control/perturbed condition. It encodes control sets with a permutation-invariant set encoder, perturbations with a structured encoder, and blends an empirical memory with a neural residual to output population mean and gene-wise variance. Trained with cross-view supervision (no cell correspondences).

Drivingauto-summary