Pub-AI: AI in Science

Driving · 2026-09-30

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei

arXiv:2609.38169PDFCode

Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.

TL;DR

STEPQuant is a spatial-temporal post-training quantization method for Delta-rule linear-attention recurrent states. It allocates quantization precision using lifetime-aware bit allocation (memory error persistence) and key-row-aware dual-axis fitting (key-row impact and row/column magnitude structure). On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, it matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 at 4-bit, with large serving memory reductions.

Why it matters

Recurrent linear attention keeps a fixed-size state, but each concurrent request needs its own persistent state, so serving memory can become a bottleneck. The authors show uniform low-precision quantization causes accuracy drops due to error propagation through Delta-rule updates, and they introduce a structured quantization approach that preserves accuracy while reducing recurrent-state memory and serving memory in SGLang.

Method

  • Lifetime-aware Bit Allocation: mixed-precision post-training quantization that assigns precision under a fixed memory budget based on quantization error magnitude and mean log retention (memory lifetime).
  • Key-Row-Aware Dual-axis Fitting: separate key-row scales and value-column scales; row factors incorporate both row magnitude and measured key-row readout impact; column scales are fit with an impact-weighted reconstruction objective.
  • Implementation in SGLang: fused kernels that reconstruct, apply the Delta update, compute readout, and write back quantized packed recurrent states; offline calibration fixes pivot selection and layout across requests.

Limitation

STEPQuant’s lifetime weight approximates error persistence through gate decay without fully modeling the time-varying, key-dependent state transition, and therefore does not fully capture the long-term effects of quantization error. With BF16 weights at four bits, KDA loses accuracy on long-generation tasks, while Qwen generates longer outputs despite retaining near-FP32 average accuracy. Evaluation covers two models under fixed hardware/workload; gains on other architectures and dynamic serving workloads remain to be verified.

Abstract (from arXiv)

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

Related papers