JEPA
Joint-embedding predictive architectures: non-generative predictive world models. (24 papers)
2026-09-30 · Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · arXiv:2609.31657
The authors benchmark frozen DOFA and SAR-JEPA embeddings on imbalanced OpenSARShip and mitigate long-tail imbalance without backbone fine-tuning by applying minority-only oversampling in embedding space, then training a lightweight classifier head. Across both foundation models, oversampling improves Macro-F1 and accuracy vs baselines, with the largest Macro-F1 gains: DOFA+ADASYN (34.39→38.56) and SAR-JEPA+SVM-SMOTE (25.89→32.30).
JEPAauto-summary
2026-09-30 · Yutong Feng, Bowen Liao, See Kiong Ng et al. · arXiv:2609.31680
The authors evaluate one JEPA instantiation for time-series forecasting across nine backbones and eleven benchmarks. They report that JEPA’s benefit is highly inconsistent across backbones: it gives consistent gains for some architectures and consistent degradation for others, even on the same dataset. The authors find the variability holds across both temporal and spatio-temporal task families.
JEPAauto-summary
2026-09-30 · Peng Xie, Amr Alanwar · arXiv:2609.32481
The authors study how JEPA masking geometry determines what can be recovered from context. They model a mask as a linear measurement: atoms whose support lies in the hidden region fall in the null space in a wavelet basis. JEPA must predict only sufficiency, enabling shortcuts with a moving-average target encoder; removing the shortcut depends on coarse-scale content left unrecoverable and on reachable context. They test predictions in 151 pre-training runs.
JEPAauto-summary
2026-09-30 · Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu · arXiv:2609.32517
LocalProp is a training procedure that locally updates model weights by restricting gradient propagation to a single active module (e.g., one transformer block) while detaching other parts. The pipeline follows “pre-training then fine-tuning” using a local variant of I-JEPA, followed by structured pruning and short local recovery, aiming to reduce peak GPU memory and control the accuracy–memory trade-off via the number of jointly optimized blocks.
JEPAauto-summary
2026-09-30 · Aarnav Choudhary · arXiv:2609.32908
The authors study whether a JEPA-style latent proof-transition objective from one-step Lean transitions can rank kernel-validated successor states for search. On a same-theorem ranking diagnostic, JEPA has higher Top-1 than matched InfoNCE (50.18% vs 31.55%). But in fixed-budget kernel-checked search, JEPA solves fewer theorems (282.3 vs proposer ordering 308) and requires more tactic checks, indicating one-step transition ranking is insufficient long-horizon search value under this setup.
JEPAauto-summary
2026-09-30 · Idan Achituve, Lior Dikstein, Idit Diamant et al. · arXiv:2609.32921
Adaptive LeWorldModel (ALeWM) is a JEPA-style world model that learns to concentrate predictive information into compact prefixes of a wide latent embedding. It uses a sequence-conditioned capacity network to sample a prefix length during training, then regularizes masked embeddings with MixSIGReg against a Gaussian-active/zero-suffix mixture target. The learned ordering supports recursive planning with lower average planning capacity than fixed-width LeWM while achieving higher mean success rates.
JEPAauto-summary
2026-09-30 · Constantino \'Alvarez Casado, Nhi Nguyen, Mohammad Rakibur Rahman et al. · arXiv:2609.32928
Phenomenon-Graph JEPA is a joint-embedding predictive architecture for contactless cardiorespiratory sensing that uses four processed 1D streams, predicts stopped target embeddings along typed edges, and uses within-episode forward prediction without negative pairs or synthetic augmentation in the base setup. On OMuSense-23, the authors report improved label-efficiency area over matched supervised training (3.91 pp) under one configuration, but no advantage of physiological edge typing versus wrong-edge/all-pairs controls.
JEPAauto-summary
2026-09-30 · Nitin Nagesh Kulkarni, Aashwin Anand Mishra, Yin Yu et al. · arXiv:2609.33110
D-JEPA learns a geometry-only latent representation and uses swappable, lightweight physics-specific decoders conditioned on operating conditions to predict physical fields. An explicit design-recoverability objective encourages the geometry latent to preserve underlying design variables for linear recovery and design optimization. The authors report that on four 3D benchmarks it maintains/improves full-field accuracy and achieves near-perfect linear recoverability while enabling decoder transfer.
JEPAauto-summary
2026-09-30 · Wanfeng Lu, Yutong Zhang, Keyi Zhou et al. · arXiv:2609.33350
KoopCell is a Koopman-based generative framework for learning continuous single-cell population dynamics from temporally sparse, unpaired distribution snapshots. It jointly learns latent coordinates and linear latent dynamics guided by Koopman–Mori–Zwanzig theory, with convergence guarantees from a weak continuity equation. KoopCell-M adds non-Markovian memory via Markovian embedding to model branching in latent dynamics.
JEPAauto-summary
2026-09-30 · Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi et al. · arXiv:2609.33487
The authors run a controlled sweep over spatio-temporal EEG masking geometry for 58 masked-prediction foundation models trained with MAE and JEPA, then evaluate all models on 12 OpenEEGBench datasets using linear probing. Both frameworks agree on an optimal mask configuration and shared failure modes; performance is robust outside the best region, while JEPA has an additional failure mode (“bias-inflation collapse”).
JEPAauto-summary
2026-09-30 · Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt · arXiv:2609.33563
MA-JEPA introduces a stochastic joint-embedding predictive world model for multi-agent RL with centralized training and decentralized execution. It replaces observation reconstruction with prediction of target observation embeddings, using a categorical latent state and a causal Transformer. A training-only joint predictor conditions on all agents’ local states/actions to predict each agent’s next embedding, which is passed through the same local posterior as in execution for actor-critic learning from latent imagination.
JEPAauto-summary
2026-09-30 · Thomas Walker, Randall Balestriero, Richard Baraniuk · arXiv:2609.33940
The authors propose behavioral monitoring signals for JEPA world models using centroids defined as sub-component Jacobian row-sums. Centroids can be computed via Jacobian vector products, characterize internal input-space geometry, and yield saliency maps. On Push-T and TwoRoom under distribution shift, a structural dissociation (encoder goal represented, predictor unresponsive) predicts planning failure before actions; centroid-based methods outperform activation- and reconstruction-based shift detectors, enabling pre-execution goal resampling.
JEPAauto-summary
2026-09-30 · Zezhong Ding, Yipeng Li, Xike Xie · arXiv:2609.34159
WorldGraph studies graph world modeling (GWM) where the evolving graph itself is the world dynamics. It formulates latent graph states and heterogeneous transition prediction over node-, edge-, and graph-level changes, builds the GWM-Zero benchmark (8 temporal graph datasets), and proposes WorldGraph with a state-aware graph transformer and a transition-aware GRPO (dynamic grouping, structure-aware verifiable rewards).
JEPAauto-summary
2026-09-30 · Luzhe Huang, Lei Chu, Jingyi Liang et al. · arXiv:2609.34375
LRC-JEPA is a compact JEPA world model that routes information into two streams: a predictive latent z for action-conditioned dynamics and planning, and a residual context embedding u for reconstruction. Only z is rolled out by the dynamics model at test time. The authors report improved planning success in simulated control and better Bridge-v2 offline action-recovery while using a 5.5M active-parameter encoder and faster planning.
JEPAauto-summary
2026-09-30 · Shidu Ren, Qilin Gu, Zhenghao Ni et al. · arXiv:2609.35138
FlexiWorld is a JEPA-based latent world model for goal-directed planning that learns from mixed-span goal supervision while using variable-length action chunks across multiple time scales. It jointly trains a flexible (variable-length) causal action encoder, latent predictor, and an autoregressive goal-conditioned actor with Student Forcing. For planning, it proposes ARCEM (actor-residual CEM), combining action-residual search with within-chunk feedback and chunk-boundary latent prediction.
JEPAauto-summary
2026-09-30 · Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca et al. · arXiv:2609.35288
The authors introduce SACReg, a spectral anti-collapse regularizer meant to prevent dimensional collapse in the backbone representation (not just the projected space) by using a log-determinant penalty on representation covariance (and a mean penalty). They derive the idea from a λ-balance analysis in a two-layer linear network, extend it to nonlinear encoders, and apply it to JEPA as λ-JEPA with SACReg on both backbone and projector. They report improved classification/transfer and video SSL results, with released code.
JEPAcodeauto-summary
2026-09-30 · Carlos Garrido-Munoz, Jorge Calvo-Zaragoza · arXiv:2609.35473
For Handwritten Text Recognition, the authors report discriminative signal concentrated in high-variance pixel directions. They test six self-supervised learning (SSL) methods across six handwriting benchmarks and find pixel-grounded masked image modeling (MAE, SimMIM) achieves the lowest CER in frozen and fine-tuned settings, and benefits from real-handwriting pretraining; JEPA-style and contrastive methods do not.
JEPAauto-summary
2026-09-30 · Antonio Pariente, Ignacio Boero, Nikolai Matni et al. · arXiv:2609.36305
The authors propose a JEPA-style world model where latent dynamics are restricted to a bilinear (then normalized) parameterization, enabling efficient gradient-based planning and structurally enforcing action recoverability to prevent representation collapse. They report that, on standard 2D/3D control tasks, bilinear-parameterized representations match or improve success while reducing planning time by nearly three orders of magnitude.
JEPAauto-summary
2026-09-30 · Jingnan Pu, Zi-En Fan, Feng Lian · arXiv:2609.36952
ER-JEPA adds an episodic replay pathway to LLM-JEPA. It stores training token pairs in a memory and, at each step, retrieves relevant stored examples to provide additional supervision for both token prediction and representation alignment. The replay path is removed after training, so inference matches LLM-JEPA.
JEPAauto-summary
2026-09-30 · Mingu Kang, Yoori Oh, Sookyung Kim et al. · arXiv:2609.37441
The authors show that joint training of JEPA-style latent world models with isotropic Gaussian regularization can learn a representation geometry whose Euclidean latent planning cost ranks feasible outcomes differently from the task cost, even with accurate prediction. They propose AnisoWM with ΛReg: replacing the fixed isotropic target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. Across four visual goal-planning environments, it improves planning success over LeWorldModel in all four.
JEPAauto-summary