Pub-AI: AI in Science

JEPA · 2026-09-30

$\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

Berker Demirel, Cl\'ementine Domin\'e, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello

arXiv:2609.35288PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

The authors introduce SACReg, a spectral anti-collapse regularizer meant to prevent dimensional collapse in the backbone representation (not just the projected space) by using a log-determinant penalty on representation covariance (and a mean penalty). They derive the idea from a λ-balance analysis in a two-layer linear network, extend it to nonlinear encoders, and apply it to JEPA as λ-JEPA with SACReg on both backbone and projector. They report improved classification/transfer and video SSL results, with released code.

Why it matters

Many JE-SSL methods regularize collapse after a projection head, while downstream uses the backbone. The authors argue this mismatch can leave the backbone low-rank, potentially limiting transfer. Their approach targets backbone spectral properties (covariance eigenvalues) via SACReg, and they evaluate it in both image and video JEPA settings with reported gains over related methods.

Method

  • Derive an anti-collapse mechanism from λ-balance in a two-layer linear network: negative layer balance yields non-collapse (full row rank) and motivates an encoder-side spectral regularizer.
  • Extend to nonlinear encoders with Nonlinear SACReg: apply a regularizer on backbone representation mean and centered covariance, including a negative log-det term to penalize small covariance eigenvalues.
  • Apply SACReg to JE-SSL via λ-JEPA: use view-averaged backbone representations and apply SACReg to both backbone and projected representations; on video, use a λ-JEPA training setup derived from LeVJEPA pipelines.
Abstract (from arXiv)

Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $\lambda$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $\lambda$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $\lambda$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose $\lambda$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $\lambda$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.

Related papers