NVIDIA's SparDA Architecture Boosts Decoding Speed 1.7x for Long-Context LLMs
HankYeomans · x · 2026-08-13
NVIDIA researchers introduced SparDA (Sparse Decoupled Attention), a novel Transformer architecture designed to overcome bottlenecks in long-context LLM inference.
Key Improvements & Results:
- Achieves 1.7x faster decoding and improves long-reasoning accuracy by 6.5 points.
- Traditional sparse attention reduces compute but still suffers from KV cache growing with sequence length (causing PCIe transfer bottlenecks) and retains $O(T^2)$ complexity in the sparse selection step.
- SparDA introduces a fourth per-layer projection called Forecast alongside Q, K, and V. It predicts the KV blocks needed by the next layer, enabling lookahead selection that overlaps CPU-to-GPU prefetch with current-layer execution.
Engineering Details:
- Decoupled from the attention query, the GQA implementation uses one Forecast head per group, reducing selection overhead.
- Adds <0.5% parameters and requires training only the Forecast projections to match the original selector's attention distribution.
More from Infra
- Sony, TSMC Deal Brings Japan Chipmaking Investment to $37bn — pstAsiatech · 2026-08-13
- YMTC Overtakes Kioxia in Flash Shipments Amid AI Boom — pstAsiatech · 2026-08-13
- Robots Now Autonomously Swapping Failed Drives in Data Centers — chris_j_paxton · 2026-08-13
- China AI Chip Networking Firm Kiwimoore Targets HK IPO at $2B Valuation — pstAsiatech · 2026-08-13
- RTX PRO 6000 Hits 2900 tok/s Running 30B Model Locally — HankYeomans · 2026-08-13
- Pichai Predicts TPUs in Space by 2027, Powered by Orbital Solar — rohanpaul_ai · 2026-08-13