Quantization Breaks RNN Memory: Error Feedback Restores GRU/LSTM Accuracy Without Retraining
Ismail Erbas · hf · 2026-09-07
This work identifies a concrete failure mode in quantized recurrent inference: state write-back rules suppress small updates, degrading temporal memory of GRUs and LSTMs at low precision.
The authors propose error feedback and residual memory mechanisms that compensate for the suppressed updates, restoring accuracy without any retraining across GRU and LSTM architectures—a practical reference for on-device recurrent deployment.
More from Infra
- It's 100,000 GPUs, not NVL72 racks: viral cluster-size claim corrected — firstadopter · 2026-09-07
- How compute-efficient is Astra? Estimates suggest 10x gap vs smaller labs — teortaxesTex · 2026-09-07
- Run Qwen3 27B Free on Kaggle: ~20 Hours of GPU Usage You Can Point Hermes At — TheMoonMidas · 2026-09-07
- Tuning draft acceptance for Qwen3.6-35B MTP speculative decoding in llama.cpp — Bulky-Priority6824 · 2026-09-07
- Google's MaxKernel: Multi-Agent System Writes TPU Kernels at Expert Level — google · 2026-09-07
- Cerebras: Layer Dropout Speeds Up LLM Training and Enables Early-Exit Inference — cerebras · 2026-09-07