Pub-AI: AI in Science

Digest

Reset

126 papers

Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling

2026-09-30 · Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · arXiv:2609.31657

The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).

JEPAauto-summary

Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?

2026-09-30 · Yutong Feng, Bowen Liao, See Kiong Ng et al. · arXiv:2609.31680

The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.

JEPAauto-summary

CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense

2026-09-30 · Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan et al. · arXiv:2609.31893

CyberWorld is a Dreamer-style world model for autonomous cyber defense that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of a defended network. Using CyberWheel, the graph-based variant reaches the deploy_then_stop control after 3.6k–15.8k environment steps (vs model-free PPO needing 2.3M–3.1M steps or failing within a 3.2M budget).

Drivingauto-summary

Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

2026-09-30 · Jiaming Zhang, Wu Yang, Shuai Tao et al. · arXiv:2609.31938

Cache-aware Conv3D lowering for embedded generative video VAEs: supported causal Conv3D calls are expressed as batched spatial Conv2D while preserving pretrained weights, causal-cache semantics, convolution parameters, bias placement, and output layout. On 64-GB Jetson AGX Orin in the Cosmos3-Edge image-to-video pipeline, the authors report ~7.32× VAE-decoder speedup and 2.21× end-to-end speedup at 25 frames.

Video World Modelsauto-summary

Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement

2026-09-30 · Hitoshi Iyatomi · arXiv:2609.32268

The authors study Self-Reconstruction Dynamics (SRD), produced by repeatedly applying a frozen autoencoder to its own reconstruction, which forms transient image/latent trajectories. Although repeated self-reconstruction degrades fidelity, SRD encodes sample-specific correction information. They propose SRD-guided Reconstruction Refinement (SRD-RR), predicting a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. Across six datasets, SRD-RR recovers 38.6% (one) to 45.3% (two) of the empirically recoverable MSE gap.

Latent Dynamicsauto-summary

Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality

2026-09-30 · Linhao Wang, Yiyan Fan, Dongjin Huang · arXiv:2609.32322

The authors show that two world models with similar total prediction error can yield very different planning outcomes when their errors occur on different state dimensions. They introduce Decision-Relevant Prediction Error (DRPE), measuring multi-step error only on decision-relevant state dimensions, and an iso-error protocol that varies error allocation while keeping total error fixed.

Latent Dynamicsauto-summary

JEPA Learns What the Mask Leaves Unrecoverable

2026-09-30 · Peng Xie, Amr Alanwar · arXiv:2609.32481

The authors study how JEPA masking geometry determines what can be recovered from context. They model a mask as a linear measurement: atoms whose support lies in the hidden region fall in the null space in a wavelet basis. JEPA must predict only sufficiency, enabling shortcuts with a moving-average target encoder; removing the shortcut depends on coarse-scale content left unrecoverable and on reachable context. They test predictions in 151 pre-training runs.

JEPAauto-summary

What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response

2026-09-30 · Enzo Nicol\'as Spotorno, Josafat Leal Filho, Ant\^onio Augusto Fr\"ohlich · arXiv:2609.32512

The authors present a measurement protocol for action-conditioned latent vehicle predictors with a physical readout. It separately tests retention of physical quantities, organization in latent space, one-step forecasting, and local response to small command perturbations via three matched response paths. In a case study on IPG CarMaker data, representations retain planar outputs, improved future-command inputs help 1s forecasts, but local command-response can diverge in latent space and cause regret in nearby-command ranking.

Latent Dynamicsauto-summary

LocalProp: Neuro-Localized Memory-Efficient Backpropagation

2026-09-30 · Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · arXiv:2609.32517

LocalProp is a training procedure that locally updates model weights by restricting gradient propagation to a single active module (e.g., one transformer block) while detaching other parts. The pipeline follows “pre-training then fine-tuning” using a local variant of I-JEPA, followed by structured pruning and short local recovery, aiming to reduce peak GPU memory and control the accuracy–memory trade-off via the number of jointly optimized blocks.

JEPAauto-summary

Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC

2026-09-30 · Yi Xian Goh, Sze Jue Yang, Hao Luan · arXiv:2609.32591

Fast-TD-MPC reduces per-step computation in data-driven MPC by adaptively switching between a fast amortized policy (System 1) and the original MPPI planner (System 2) using an OOD gate in TD-MPC2’s latent space. Across 103 continuous control tasks, it achieves up to ~4× faster inference, with robustness under external disturbances comparable to TD-MPC2 when planning is triggered selectively.

Drivingauto-summary

The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models

2026-09-30 · Dongsheng Liu, Chao Jin, Wenkui Yang et al. · arXiv:2609.32679

GUI world models often predict future interfaces using only the current GUI observation and action, which can cause state aliasing: identical visible conditions correspond to different valid futures due to hidden transition-relevant state. The authors introduce StateAliasBench (strict-pair diagnostic) and a predictive-state recovery method that infers structured state from history and condition frozen GUI world models to restore state-sensitive prediction and improve AndroidWorld agent performance.

Drivingauto-summary

Copper-Policy: Focus on the Representation for Robust Robot Manipulation

2026-09-30 · Zexin Feng, Yixu Feng, Lingyu Xiao et al. · arXiv:2609.32779

Copper-Policy learns a compact, task-conditioned world representation jointly with the policy using temporal joint-embedding prediction. It predicts future observation embeddings (not pixels) and decodes actions from current-frame details plus the learned compact representation; no future is generated at test time. The authors report strong control and efficient training, including a 2B model trained in 9.67 hours on 8×RTX 5090 and faster training than Fast-WAM on matched A100 GPUs.

Drivingauto-summary

Predicting the Next State Is Not Enough: JEPA Representations for Lean Theorem Proving

2026-09-30 · Aarnav Choudhary · arXiv:2609.32908

The authors study whether a JEPA-style latent proof-transition objective from one-step Lean transitions can rank kernel-validated successor states for search. On a same-theorem ranking diagnostic, JEPA has higher Top-1 than matched InfoNCE (50.18% vs 31.55%). But in fixed-budget kernel-checked search, JEPA solves fewer theorems (282.3 vs proposer ordering 308) and requires more tactic checks, indicating one-step transition ranking is insufficient long-horizon search value under this setup.

JEPAauto-summary

Adaptive Latent Capacity for World Models

2026-09-30 · Idan Achituve, Lior Dikstein, Idit Diamant et al. · arXiv:2609.32921

Adaptive LeWorldModel (ALeWM) is a JEPA-style world model that learns to concentrate predictive information into compact prefixes of a wide latent embedding. It uses a sequence-conditioned capacity network to sample a prefix length during training, then regularizes masked embeddings with MixSIGReg against a Gaussian-active/zero-suffix mixture target. The learned ordering supports recursive planning with lower average planning capacity than fixed-width LeWM while achieving higher mean success rates.

JEPAauto-summary

Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing

2026-09-30 · Constantino \'Alvarez Casado, Nhi Nguyen, Mohammad Rakibur Rahman et al. · arXiv:2609.32928

Phenomenon-Graph JEPA is a joint-embedding predictive architecture for contactless cardiorespiratory sensing that uses four processed 1D streams, predicts stopped target embeddings along typed edges, and uses within-episode forward prediction without negative pairs or synthetic augmentation in the base setup. On OMuSense-23, the authors report improved label-efficiency area over matched supervised training (3.91 pp) under one configuration, but no advantage of physiological edge typing versus wrong-edge/all-pairs controls.

JEPAauto-summary

Self-Confirming Superposition Traps in Reinforcement Learning

2026-09-30 · Dai Shi, Andi Han, Feng Chen et al. · arXiv:2609.32966

The authors show that RL’s representation-fitting loop can sustain a lower-return policy: in a self-confirming superposition trap, globally optimal codes under the current policy share overlapping directions for features that rarely co-activate, yet interfere after an alternative action. Interventions that preserve access to neglected states or protect replay weights reduce interference and can improve control, including in DreamerV3–Crafter.

Latent Dynamicsauto-summary

What Must a World Model Distinguish for Planning?

2026-09-30 · Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu et al. · arXiv:2609.33030

The authors formalize when a world model must preserve physical distinctions for planning, via mechanism/response/decision sufficiency and query-dependent resolution requirements. Experiments on collision, learned nonlinear dynamics, and robotic planning show that needed information depends on the query, candidate set, and planning stage. They also compare query placement: a modular design (query-aware proposal + query-independent action-conditioned prediction) reuses physical predictions across objectives and achieves lower held-out regret than action-conditioned prediction methods.

Drivingauto-summary

D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders

2026-09-30 · Nitin Nagesh Kulkarni, Aashwin Anand Mishra, Yin Yu et al. · arXiv:2609.33110

D-JEPA learns a geometry-only latent representation and uses swappable, lightweight physics-specific decoders conditioned on operating conditions to predict physical fields. An explicit design-recoverability objective encourages the geometry latent to preserve underlying design variables for linear recovery and design optimization. The authors report that on four 3D benchmarks it maintains/improves full-field accuracy and achieves near-perfect linear recoverability while enabling decoder transfer.

JEPAauto-summary

Dynamic Manipulation with World-Action Models via Counterfactual Planning

2026-09-30 · Sunwoo Park, Wonbin Lee, Seonghyun Jin et al. · arXiv:2609.33172

World–Action models trained on static demonstrations can fail on moving targets due to “target-response collapse” during reactive replanning. The authors propose Dynamic Predictive Planning (DPP), which predicts interaction timing using WAM rollouts, forecasts the target’s interaction position, synthesizes a counterfactual observation in a familiar robot context, and connects the resulting canonical plan to the robot’s live execution. DPP improves dynamic manipulation in simulation and on a real robot without additional training on dynamic data.

Drivingauto-summary

SMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEs

2026-09-30 · Yangyuan Li, Weichao Li, Shaowu Pan · arXiv:2609.33205

SMORE is a mesh-agnostic reduced-order model for time-dependent PDEs. It uses an INR autodecoder to map sparse spatial measurements to a latent state, then evolves the latent state with structured latent dynamics (linear or linear-quadratic). A Lyapunov-guided stability regularization is added to promote stable long-horizon rollouts, and the authors provide theoretical guarantees under structural assumptions.

Latent Dynamicsauto-summary