Liquid AI's QAD technique lets Q4 models retain 97% of BF16 performance

alexcovo_eth · x · 2026-08-20

Liquid AI released Quantization-Aware Distillation (QAD), a method where the model trains while experiencing 4-bit quantization errors, guided by a high-precision teacher model. This allows the quantized model to compensate for precision loss. Applied to LFM2.5 models (230M-2.6B), the resulting Q40 GGUFs retain roughly 97% of their BF16 performance while maintaining a small footprint.

Related event: Liquid AI Unveils Quantization-Aware Distillation, 4-bit Models Keep 97% Performance(3 posts)→

Original post →

More from Research

Research channel →