Pub-AI: AI in Science

JEPA · 2026-09-30

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

Gaspard Bott\'e, S\'everin Baroudi, Samir Sadok, Francesco Paissan, Thomas Hueber, Xavier Alameda-Pineda, Ricard Marxer, Mirco Ravanelli

arXiv:2609.37798PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

GLaS-JEPA is a speech SSL framework that predicts the current encoder’s continuous representations at masked positions, without contrastive learning, discrete targets, or a separate EMA target encoder. It prevents representation collapse using SIGReg representation-space regularization. Pretrained on 960h LibriSpeech, the 57M model achieves 6.89% WER on frozen-encoder SUPERB ASR and 25.87% CER on slot filling, outperforming listed non-distilled sub-90M baselines.

Why it matters

The authors address the question of whether engineered prediction targets are necessary for competitive speech SSL. Their approach removes discrete target generation and EMA target encoders by using current-encoder continuous targets with SIGReg regularization. The reported SUPERB transfer results aim to show that a simplified training recipe can still produce strong speech representations across content, speaker, and semantic tasks.

Method

  • Architecture: shared encoder with two paths (full view and masked view) and a token-wise linear projector mapping encoder outputs into a 128D loss space; no separate temporal predictor or EMA teacher.
  • Training objective: minimize masked-view MSE between predicted latents and stop-gradient targets from the full-view representations at masked positions.
  • Collapse prevention: SIGReg (Full Marginal in the main model) regularizes unmasked full-view target representations toward an isotropic Gaussian reference distribution in representation space.

Limitation

The authors note that published baseline comparisons use different architectures and training budgets, so they cannot isolate the SSL objective’s effect; these comparisons test whether such a target-free model can be competitive rather than whether its objective outperforms alternatives under controlled conditions.

Abstract (from arXiv)

Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.

Related papers