2026-09-30 · Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · arXiv:2609.31657
The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).
2026-09-30 · Yutong Feng, Bowen Liao, See Kiong Ng et al. · arXiv:2609.31680
The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.
CyberWorld is a Dreamer-style world model for autonomous cyber defense that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of a defended network. Using CyberWheel, the graph-based variant reaches the deploy_then_stop control after 3.6k–15.8k environment steps (vs model-free PPO needing 2.3M–3.1M steps or failing within a 3.2M budget).
2026-09-30 · Jiaming Zhang, Wu Yang, Shuai Tao et al. · arXiv:2609.31938
Cache-aware Conv3D lowering for embedded generative video VAEs: supported causal Conv3D calls are expressed as batched spatial Conv2D while preserving pretrained weights, causal-cache semantics, convolution parameters, bias placement, and output layout. On 64-GB Jetson AGX Orin in the Cosmos3-Edge image-to-video pipeline, the authors report ~7.32× VAE-decoder speedup and 2.21× end-to-end speedup at 25 frames.
The authors study Self-Reconstruction Dynamics (SRD), produced by repeatedly applying a frozen autoencoder to its own reconstruction, which forms transient image/latent trajectories. Although repeated self-reconstruction degrades fidelity, SRD encodes sample-specific correction information. They propose SRD-guided Reconstruction Refinement (SRD-RR), predicting a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. Across six datasets, SRD-RR recovers 38.6% (one) to 45.3% (two) of the empirically recoverable MSE gap.
The authors show that two world models with similar total prediction error can yield very different planning outcomes when their errors occur on different state dimensions. They introduce Decision-Relevant Prediction Error (DRPE), measuring multi-step error only on decision-relevant state dimensions, and an iso-error protocol that varies error allocation while keeping total error fixed.
2026-09-30 · Peng Xie, Amr Alanwar · arXiv:2609.32481
The authors study how JEPA masking geometry determines what can be recovered from context. They model a mask as a linear measurement: atoms whose support lies in the hidden region fall in the null space in a wavelet basis. JEPA must predict only sufficiency, enabling shortcuts with a moving-average target encoder; removing the shortcut depends on coarse-scale content left unrecoverable and on reachable context. They test predictions in 151 pre-training runs.
The authors present a measurement protocol for action-conditioned latent vehicle predictors with a physical readout. It separately tests retention of physical quantities, organization in latent space, one-step forecasting, and local response to small command perturbations via three matched response paths. In a case study on IPG CarMaker data, representations retain planar outputs, improved future-command inputs help 1s forecasts, but local command-response can diverge in latent space and cause regret in nearby-command ranking.
LocalProp is a training procedure that locally updates model weights by restricting gradient propagation to a single active module (e.g., one transformer block) while detaching other parts. The pipeline follows “pre-training then fine-tuning” using a local variant of I-JEPA, followed by structured pruning and short local recovery, aiming to reduce peak GPU memory and control the accuracy–memory trade-off via the number of jointly optimized blocks.
Fast-TD-MPC reduces per-step computation in data-driven MPC by adaptively switching between a fast amortized policy (System 1) and the original MPPI planner (System 2) using an OOD gate in TD-MPC2’s latent space. Across 103 continuous control tasks, it achieves up to ~4× faster inference, with robustness under external disturbances comparable to TD-MPC2 when planning is triggered selectively.
2026-09-30 · Dongsheng Liu, Chao Jin, Wenkui Yang et al. · arXiv:2609.32679
GUI world models often predict future interfaces using only the current GUI observation and action, which can cause state aliasing: identical visible conditions correspond to different valid futures due to hidden transition-relevant state. The authors introduce StateAliasBench (strict-pair diagnostic) and a predictive-state recovery method that infers structured state from history and condition frozen GUI world models to restore state-sensitive prediction and improve AndroidWorld agent performance.
Copper-Policy learns a compact, task-conditioned world representation jointly with the policy using temporal joint-embedding prediction. It predicts future observation embeddings (not pixels) and decodes actions from current-frame details plus the learned compact representation; no future is generated at test time. The authors report strong control and efficient training, including a 2B model trained in 9.67 hours on 8×RTX 5090 and faster training than Fast-WAM on matched A100 GPUs.
The authors study whether a JEPA-style latent proof-transition objective from one-step Lean transitions can rank kernel-validated successor states for search. On a same-theorem ranking diagnostic, JEPA has higher Top-1 than matched InfoNCE (50.18% vs 31.55%). But in fixed-budget kernel-checked search, JEPA solves fewer theorems (282.3 vs proposer ordering 308) and requires more tactic checks, indicating one-step transition ranking is insufficient long-horizon search value under this setup.
Adaptive LeWorldModel (ALeWM) is a JEPA-style world model that learns to concentrate predictive information into compact prefixes of a wide latent embedding. It uses a sequence-conditioned capacity network to sample a prefix length during training, then regularizes masked embeddings with MixSIGReg against a Gaussian-active/zero-suffix mixture target. The learned ordering supports recursive planning with lower average planning capacity than fixed-width LeWM while achieving higher mean success rates.
2026-09-30 · Constantino \'Alvarez Casado, Nhi Nguyen, Mohammad Rakibur Rahman et al. · arXiv:2609.32928
Phenomenon-Graph JEPA is a joint-embedding predictive architecture for contactless cardiorespiratory sensing that uses four processed 1D streams, predicts stopped target embeddings along typed edges, and uses within-episode forward prediction without negative pairs or synthetic augmentation in the base setup. On OMuSense-23, the authors report improved label-efficiency area over matched supervised training (3.91 pp) under one configuration, but no advantage of physiological edge typing versus wrong-edge/all-pairs controls.
2026-09-30 · Dai Shi, Andi Han, Feng Chen et al. · arXiv:2609.32966
The authors show that RL’s representation-fitting loop can sustain a lower-return policy: in a self-confirming superposition trap, globally optimal codes under the current policy share overlapping directions for features that rarely co-activate, yet interfere after an alternative action. Interventions that preserve access to neglected states or protect replay weights reduce interference and can improve control, including in DreamerV3–Crafter.
2026-09-30 · Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu et al. · arXiv:2609.33030
The authors formalize when a world model must preserve physical distinctions for planning, via mechanism/response/decision sufficiency and query-dependent resolution requirements. Experiments on collision, learned nonlinear dynamics, and robotic planning show that needed information depends on the query, candidate set, and planning stage. They also compare query placement: a modular design (query-aware proposal + query-independent action-conditioned prediction) reuses physical predictions across objectives and achieves lower held-out regret than action-conditioned prediction methods.
D-JEPA learns a geometry-only latent representation and uses swappable, lightweight physics-specific decoders conditioned on operating conditions to predict physical fields. An explicit design-recoverability objective encourages the geometry latent to preserve underlying design variables for linear recovery and design optimization. The authors report that on four 3D benchmarks it maintains/improves full-field accuracy and achieves near-perfect linear recoverability while enabling decoder transfer.
2026-09-30 · Sunwoo Park, Wonbin Lee, Seonghyun Jin et al. · arXiv:2609.33172
World–Action models trained on static demonstrations can fail on moving targets due to “target-response collapse” during reactive replanning. The authors propose Dynamic Predictive Planning (DPP), which predicts interaction timing using WAM rollouts, forecasts the target’s interaction position, synthesizes a counterfactual observation in a familiar robot context, and connects the resulting canonical plan to the robot’s live execution. DPP improves dynamic manipulation in simulation and on a real robot without additional training on dynamic data.
2026-09-30 · Yangyuan Li, Weichao Li, Shaowu Pan · arXiv:2609.33205
SMORE is a mesh-agnostic reduced-order model for time-dependent PDEs. It uses an INR autodecoder to map sparse spatial measurements to a latent state, then evolves the latent state with structured latent dynamics (linear or linear-quadratic). A Lyapunov-guided stability regularization is added to promote stable long-horizon rollouts, and the authors provide theoretical guarantees under structural assumptions.