USC and Yale propose KronQ for state-of-the-art 2-bit LLaMA-3-70B quantization
burkov · x · 2026-07-24
Researchers from USC and Yale introduce KronQ, a post-training quantization framework for large language models.
What it does
- Targets 2-bit weight-only quantization on LLaMA-3-70B
- Incorporates gradient covariance via a Kronecker-factored Hessian
- The authors say it reaches state-of-the-art results for this setting
Why it matters
- Existing methods reportedly fail to converge in this regime
- KronQ is positioned as a more stable way to push LLM compression to very low bit-widths without collapsing optimization
The post links to an AI tutor for reading the work.
More from Research
- Composite-Bench debuts with verified computer-use evals; GLM-5.2 leads Kimi K3 by 32 points — davidtsong · 2026-07-24
- Arsenal hires a research engineer to build AI models for football analysis — Vjeux · 2026-07-24
- GEPA composes optimizers into meta-optimizer pipelines in an "optimize_anything" framework — Thrumpwart · 2026-07-24
- NeurIPS meta-review-before-rebuttal process gets challenged as anchoring decisions too early — xwang_lk · 2026-07-24
- Embodied AI needs years of expert trade knowledge to learn real-world constraints — Exp_Mark · 2026-07-24
- Custom eval harness ranks Fable 5 and Opus 4.6 above 10 models — rudrank · 2026-07-24