Pub-AI: AI in Science

JEPA · 2026-09-30

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Carlos Garrido-Munoz, Jorge Calvo-Zaragoza

arXiv:2609.35473PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

For Handwritten Text Recognition, the authors report discriminative signal concentrated in high-variance pixel directions. They test six self-supervised learning (SSL) methods across six handwriting benchmarks and find pixel-grounded masked image modeling (MAE, SimMIM) achieves the lowest CER in frozen and fine-tuned settings, and benefits from real-handwriting pretraining; JEPA-style and contrastive methods do not.

Why it matters

The paper argues that SSL objectives should match where the task’s discriminative information resides in input space. For HTR, pixel-space reconstruction is not wasted capacity: preserving high-variance pixel content predicts better character error rates, and a frozen pixel-grounded encoder can be competitive with fully fine-tuned supervised baselines when paired with a pretrained LLM decoder.

Method

  • Test six SSL methods (MAE, SimMIM, I-JEPA, V-JEPA-2, SigLIP, MoCo-v3) under matched encoder/data/evaluation protocols on six line-level HTR benchmarks in five languages; use CER with frozen probes and LLM-based multi-stage fine-tuning.
  • Use a projection-then-probe protocol: compute pixel PCA and compare top-variance vs bottom-variance reconstructions, then train the same BiLSTM-CTC probe to measure which subspace preserves text recognizability.
  • Analyze how encoder features align with the high-variance pixel subspace and relate that alignment to CER across methods; also evaluate label efficiency with limited supervision.
Abstract (from arXiv)

In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.

Related papers