TQ: calibration-free 4-bit quantization open-sourced, hits 92.4% top-1 on Qwen 27B
textclf · reddit · 2026-09-28
A developer open-sourced TQ, a calibration-free quantization method, together with its implementation 'Quant Factory':
- Key claim: needs no calibration data, so models can be quantized the moment they ship, and quantized weights generalize better while staying close to calibration-based methods
- Numbers: 4-bit TQ on Qwen 3.8 27B achieves mean KLD of 0.0282 and 92.4% top-1
- Method: grounded in information-theoretic lossy source coding, claimed optimal at the rate-distortion limit
- Practicality: 4-bit only for now (2-bit/3-bit planned); quantized models published on Hugging Face, with a Docker image that serves them via vLLM using --quantization tq
More from Infra
- Warehouse full of servers sells private isolated AI model hosting per client — rohanpaul_ai · 2026-09-28
- Is Agentic scores how AI-agent-ready your website is, via a single npx command — seanwbren · 2026-09-28
- Estimating 100M DAU infra for Meta Muse: 1-4GW of power, tiny $3B sandbox layer — SuB8u · 2026-09-28
- $350 Dell from 2007 beats $1500 RTX 5070 rig on agentic LLM tasks — Truth-Does-Not-Exist · 2026-09-28
- Rural town promises every household $10k if a data center gets built — JumpCrisscross · 2026-09-28
- GPU-backed loans cost ~1.2pts over normal loans — lenders only trust assets that outlive chip generations — rohanpaul_ai · 2026-09-28