Paper: CKA-QAD Significantly Improves Reasoning and Coding in NVFP4 Quantized LLMs
Aaaaaaaaaeeeee · reddit · 2026-08-10
The arXiv paper Beyond Output Matching delves into the accuracy recovery problem of large language models under low-bit quantization (e.g., NVFP4).
- The Problem: Traditional Quantization-Aware Distillation (QAD) relies solely on KL-divergence to match output distributions, which can mask internal representational degradation. This causes severe drift in intermediate activation geometries, especially in RL-post-trained models, leading to bottlenecks in reasoning and coding tasks.
- The Solution: The authors propose CKA-QAD, a CKA (Centered Kernel Alignment)-guided representational alignment method. It preserves internal representational geometry during distillation by aligning layerwise Gram matrices via a lightweight regularizer.
- Results: Experiments on Nemotron 3 Nano and Qwen3-4B-Thinking-2507 show that CKA-QAD substantially improves representational alignment and boosts downstream reasoning and coding accuracy with modest training overhead.
More from Infra
- NVIDIA Shares DGX Spark Guide for Local LLM Deployment — lifebypixels · 2026-08-10
- AGI Bottlenecks: Infrastructure Costs and High-Bandwidth Comms — ns123abc · 2026-08-10
- Data Center Tax Boom: Small ND Town Sees 7x Property Tax Base Surge — BenBajarin · 2026-08-10
- Musk Predicts AI Agent Internet Traffic Will Vastly Exceed Human Usage — elonmusk · 2026-08-10
- AI Hedge Fund Invests $400M in Chip Startup Source Foundry — TechCrunch AI · 2026-08-10
- Google Open-Sources $80 Offline AI Translator Running Gemma on Raspberry Pi — alexanderchen · 2026-08-10