Disaggregated Quantization Boosts 1-bit LLM Accuracy by 32+ Points and TTFT by 1.78x
ISTA-DASLab · hf · 2026-09-29
ISTA-DASLab proposes Disaggregated Quantization (DQ), which specializes compute formats, weights, and storage placement separately for prefill and decode phases. Highlights:
- On Qwen 3 and Gemma 3, removing activation quantization only on decode improves decode-heavy task accuracy at no extra inference cost
- Training compute-native prefill weights matches or beats weight-only inference at 2-3-bit decode
- Training an NVFP4 prefiller on top of released Qwen3.8-27B GGUF decoders lifts 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without touching the decode checkpoint
- Offloaded disaggregated prefill (ODP) streams weights from SSD, delivering a 1.78x time-to-first-token speedup at 8K prompts in llama.cpp
- Validated in vLLM disaggregated serving and via post-training quantization on models up to 2.8T parameters
More from Infra
- Only 3 of ~6,000 data center projects hit by AI buildout moratoriums: SemiAnalysis — MatthewBerman · 2026-09-29
- Cloudflare birthday week ships 8 open source updates: forge, vinext 1.0, native Rust in Workers — ritakozlov · 2026-09-29
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- Sentdex has run 4B+ tokens locally on GLM 5.3 Flash — his most-used local model ever — Sentdex · 2026-09-29
- Kipply breaks down transformer inference arithmetic for H200/B200 in new perf engineering repo — ycombinator · 2026-09-29
- Dev asks Astra to 'make my GPUs not as hot' — and it apparently works — TheZachMueller · 2026-09-29