Pub-AI: AI in Science

JEPA · 2026-09-30

LocalProp: Neuro-Localized Memory-Efficient Backpropagation

Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu

arXiv:2609.32517PDF

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

LocalProp is a training procedure that locally updates model weights by restricting gradient propagation to a single active module (e.g., one transformer block) while detaching other parts. The pipeline follows “pre-training then fine-tuning” using a local variant of I-JEPA, followed by structured pruning and short local recovery, aiming to reduce peak GPU memory and control the accuracy–memory trade-off via the number of jointly optimized blocks.

Why it matters

The authors target a memory and dependency bottleneck of full end-to-end backpropagation in multi-stage training pipelines. They propose unifying local I-JEPA pre-training, local supervised fine-tuning, structured pruning, and local recovery under a constrained gradient budget, and report accuracy and peak-memory trade-offs plus comparisons to full backpropagation and other non-full-backprop methods.

Method

  • LocalProp restricts optimization to one active block/group at a time, detaching computation at module boundaries so gradients are limited to the active module and its corresponding predictor/classifier.
  • Local I-JEPA pre-training uses a block-wise objective where the target-encoder outputs are evaluated without gradients, and only the selected block’s predictor is optimized (stop-gradient for targets; block chosen round-robin).
  • Structured pruning is scored via gradient-free ablation using downstream classification loss; after pruning, the authors enforce structured masks during a final round of block-local supervised recovery.

Limitation

The authors report misalignment with full-backpropagation: “the alignment is imperfect… indicating that LocalProp models do not receive the complete coordination signal available through full-backpropagation.”

Abstract (from arXiv)

The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly localized. Therefore, we propose LocalProp, a training procedure that locally updates the weights of a model. Our neuro-localized weight updates follow the "pre-training then fine-tuning" paradigm, where the pre-training is based on I-JEPA. After locally updating the weights, a pruning operation is performed, followed by a short final fine-tuning phase. Pruning helps by sending the learning signal from higher blocks to lower blocks. We perform experiments on several datasets, including large-scale benchmarks such as ImageNet, and empirically show that LocalProp reaches good performance at a fraction of GPU peak memory. By varying the number of jointly optimized blocks, we identify gradient-propagation span as a practical control over the accuracy-memory trade-off.

Related papers