Dev asks if HY4's 1.25-bit quantization (1.5TB→200GB, 98% retention) is worth porting to Qwen3.8-Flash-Next

TemperatureOk3561 · reddit · 2026-08-31

A Reddit user asks whether HY4's 1.25-bit quantization approach — which shrank the model from 1.5TB to 200GB with 98% benchmark retention, documented in papers — is worth applying to Qwen3.8-Flash-Next, and whether it's a feasible personal project. They admit to habitually underestimating project scope and want community feedback first.

Original post →

More from Infra

Infra channel →