Pub-AI: AI in Science

JEPA

Joint-embedding predictive architectures: non-generative predictive world models. (24 papers)

Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling

2026-09-30 · Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · arXiv:2609.31657

The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).

JEPAauto-summary

Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?

2026-09-30 · Yutong Feng, Bowen Liao, See Kiong Ng et al. · arXiv:2609.31680

The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.

JEPAauto-summary

JEPA Learns What the Mask Leaves Unrecoverable

2026-09-30 · Peng Xie, Amr Alanwar · arXiv:2609.32481

The authors study how JEPA masking geometry determines what can be recovered from context. They model a mask as a linear measurement: atoms whose support lies in the hidden region fall in the null space in a wavelet basis. JEPA must predict only sufficiency, enabling shortcuts with a moving-average target encoder; removing the shortcut depends on coarse-scale content left unrecoverable and on reachable context. They test predictions in 151 pre-training runs.

JEPAauto-summary

LocalProp: Neuro-Localized Memory-Efficient Backpropagation

2026-09-30 · Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · arXiv:2609.32517

LocalProp is a training procedure that locally updates model weights by restricting gradient propagation to a single active module (e.g., one transformer block) while detaching other parts. The pipeline follows “pre-training then fine-tuning” using a local variant of I-JEPA, followed by structured pruning and short local recovery, aiming to reduce peak GPU memory and control the accuracy–memory trade-off via the number of jointly optimized blocks.

JEPAauto-summary

Predicting the Next State Is Not Enough: JEPA Representations for Lean Theorem Proving

2026-09-30 · Aarnav Choudhary · arXiv:2609.32908

The authors study whether a JEPA-style latent proof-transition objective from one-step Lean transitions can rank kernel-validated successor states for search. On a same-theorem ranking diagnostic, JEPA has higher Top-1 than matched InfoNCE (50.18% vs 31.55%). But in fixed-budget kernel-checked search, JEPA solves fewer theorems (282.3 vs proposer ordering 308) and requires more tactic checks, indicating one-step transition ranking is insufficient long-horizon search value under this setup.

JEPAauto-summary

Adaptive Latent Capacity for World Models

2026-09-30 · Idan Achituve, Lior Dikstein, Idit Diamant et al. · arXiv:2609.32921

Adaptive LeWorldModel (ALeWM) is a JEPA-style world model that learns to concentrate predictive information into compact prefixes of a wide latent embedding. It uses a sequence-conditioned capacity network to sample a prefix length during training, then regularizes masked embeddings with MixSIGReg against a Gaussian-active/zero-suffix mixture target. The learned ordering supports recursive planning with lower average planning capacity than fixed-width LeWM while achieving higher mean success rates.

JEPAauto-summary

Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing

2026-09-30 · Constantino \'Alvarez Casado, Nhi Nguyen, Mohammad Rakibur Rahman et al. · arXiv:2609.32928

Phenomenon-Graph JEPA is a joint-embedding predictive architecture for contactless cardiorespiratory sensing that uses four processed 1D streams, predicts stopped target embeddings along typed edges, and uses within-episode forward prediction without negative pairs or synthetic augmentation in the base setup. On OMuSense-23, the authors report improved label-efficiency area over matched supervised training (3.91 pp) under one configuration, but no advantage of physiological edge typing versus wrong-edge/all-pairs controls.

JEPAauto-summary

D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders

2026-09-30 · Nitin Nagesh Kulkarni, Aashwin Anand Mishra, Yin Yu et al. · arXiv:2609.33110

D-JEPA learns a geometry-only latent representation and uses swappable, lightweight physics-specific decoders conditioned on operating conditions to predict physical fields. An explicit design-recoverability objective encourages the geometry latent to preserve underlying design variables for linear recovery and design optimization. The authors report that on four 3D benchmarks it maintains/improves full-field accuracy and achieves near-perfect linear recoverability while enabling decoder transfer.

JEPAauto-summary

KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots

2026-09-30 · Wanfeng Lu, Yutong Zhang, Keyi Zhou et al. · arXiv:2609.33350

KoopCell is a Koopman-based generative framework for learning continuous single-cell population dynamics from temporally sparse, unpaired distribution snapshots. It jointly learns latent coordinates and linear latent dynamics guided by Koopman–Mori–Zwanzig theory, with convergence guarantees from a weak continuity equation. KoopCell-M adds non-Markovian memory via Markovian embedding to model branching in latent dynamics.

JEPAauto-summary

What masking geometry works best for EEG foundation models?

2026-09-30 · Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi et al. · arXiv:2609.33487

The authors run a controlled sweep over spatio-temporal EEG masking geometry for 58 masked-prediction foundation models trained with MAE and JEPA, then evaluate all models on 12 OpenEEGBench datasets using linear probing. Both frameworks agree on an optimal mask configuration and shared failure modes; performance is robust outside the best region, while JEPA has an additional failure mode (“bias-inflation collapse”).

JEPAauto-summary

MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

2026-09-30 · Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt · arXiv:2609.33563

MA-JEPA introduces a stochastic joint-embedding predictive world model for multi-agent RL with centralized training and decentralized execution. It replaces observation reconstruction with prediction of target observation embeddings, using a categorical latent state and a causal Transformer. A training-only joint predictor conditions on all agents’ local states/actions to predict each agent’s next embedding, which is passed through the same local posterior as in execution for actor-critic learning from latent imagination.

JEPAauto-summary

Behavioral Monitoring of JEPA World Models with Jacobian Centroids

2026-09-30 · Thomas Walker, Randall Balestriero, Richard Baraniuk · arXiv:2609.33940

The authors propose behavioral monitoring signals for JEPA world models using centroids defined as sub-component Jacobian row-sums. Centroids can be computed via Jacobian vector products, characterize internal input-space geometry, and yield saliency maps. On Push-T and TwoRoom under distribution shift, a structural dissociation (encoder goal represented, predictor unresponsive) predicts planning failure before actions; centroid-based methods outperform activation- and reconstruction-based shift detectors, enabling pre-execution goal resampling.

JEPAauto-summary

WorldGraph: Graph-Native World Modeling

2026-09-30 · Zezhong Ding, Yipeng Li, Xike Xie · arXiv:2609.34159

WorldGraph studies graph world modeling (GWM) where the evolving graph itself is the world dynamics. It formulates latent graph states and heterogeneous transition prediction over node-, edge-, and graph-level changes, builds the GWM-Zero benchmark (8 temporal graph datasets), and proposes WorldGraph with a state-aware graph transformer and a transition-aware GRPO (dynamic grouping, structure-aware verifiable rewards).

JEPAauto-summary

LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models

2026-09-30 · Luzhe Huang, Lei Chu, Jingyi Liang et al. · arXiv:2609.34375

LRC-JEPA is a compact JEPA world model that routes information into two streams: a predictive latent z for action-conditioned dynamics and planning, and a residual context embedding u for reconstruction. Only z is rolled out by the dynamics model at test time. The authors report improved planning success in simulated control and better Bridge-v2 offline action-recovery while using a 5.5M active-parameter encoder and faster planning.

JEPAauto-summary

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

2026-09-30 · Shidu Ren, Qilin Gu, Zhenghao Ni et al. · arXiv:2609.35138

FlexiWorld is a JEPA-based latent world model for goal-directed planning that learns from mixed-span goal supervision while using variable-length action chunks across multiple time scales. It jointly trains a flexible (variable-length) causal action encoder, latent predictor, and an autoregressive goal-conditioned actor with Student Forcing. For planning, it proposes ARCEM (actor-residual CEM), combining action-residual search with within-chunk feedback and chunk-boundary latent prediction.

JEPAauto-summary

$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

2026-09-30 · Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca et al. · arXiv:2609.35288

The authors introduce SACReg, a spectral anti-collapse regularizer meant to prevent dimensional collapse in the backbone representation (not just the projected space) by using a log-determinant penalty on representation covariance (and a mean penalty). They derive the idea from a λ-balance analysis in a two-layer linear network, extend it to nonlinear encoders, and apply it to JEPA as λ-JEPA with SACReg on both backbone and projector. They report improved classification/transfer and video SSL results, with released code.

JEPAcodeauto-summary

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

2026-09-30 · Carlos Garrido-Munoz, Jorge Calvo-Zaragoza · arXiv:2609.35473

For Handwritten Text Recognition, the authors report discriminative signal concentrated in high-variance pixel directions. They test six self-supervised learning (SSL) methods across six handwriting benchmarks and find pixel-grounded masked image modeling (MAE, SimMIM) achieves the lowest CER in frozen and fine-tuned settings, and benefits from real-handwriting pretraining; JEPA-style and contrastive methods do not.

JEPAauto-summary

Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control

2026-09-30 · Antonio Pariente, Ignacio Boero, Nikolai Matni et al. · arXiv:2609.36305

The authors propose a JEPA-style world model where latent dynamics are restricted to a bilinear (then normalized) parameterization, enabling efficient gradient-based planning and structurally enforcing action recoverability to prevent representation collapse. They report that, on standard 2D/3D control tasks, bilinear-parameterized representations match or improve success while reducing planning time by nearly three orders of magnitude.

JEPAauto-summary

ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models

2026-09-30 · Jingnan Pu, Zi-En Fan, Feng Lian · arXiv:2609.36952

ER-JEPA adds an episodic replay pathway to LLM-JEPA. It stores training token pairs in a memory and, at each step, retrieves relevant stored examples to provide additional supervision for both token prediction and representation alignment. The replay path is removed after training, so inference matches LLM-JEPA.

JEPAauto-summary

Anisotropic Representations Improve Planning in JEPA World Models

2026-09-30 · Mingu Kang, Yoori Oh, Sookyung Kim et al. · arXiv:2609.37441

The authors show that joint training of JEPA-style latent world models with isotropic Gaussian regularization can learn a representation geometry whose Euclidean latent planning cost ranks feasible outcomes differently from the task cost, even with accurate prediction. They propose AnisoWM with ΛReg: replacing the fixed isotropic target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. Across four visual goal-planning environments, it improves planning success over LeWorldModel in all four.

JEPAauto-summary