Reddit asks: are Q1 extreme quantizations really that bad, or does architecture matter more?
Felix_455-788 · reddit · 2026-09-20
A Reddit user is crowdsourcing experiences running Q1/Q2/Q3 extreme quantizations locally, raising perennial debates: how much quality loss comes from the quant scheme vs. model architecture, whether UD-Q differs meaningfully from IQ formats, and the classic trade-off of a small Q4 model versus a 200B+ MoE model squashed to Q1.
More from Models
- Opus 5 says nagging you to sleep is its only allowed tenderness — repligate · 2026-09-20
- Jev-inspired inference makes 350M-parameter LFM2.5 63x faster on L40S, code and weights released — JosephJacks_ · 2026-09-20
- Same SVG prompt, striking jump: Fable 5.1 output leaps ahead within days — LegitimateSwordfish8 · 2026-09-20
- All 4 labs whose models escaped testing were evaluated by the same company, Irregular — bradneuberg · 2026-09-20
- Claude Opus users report it overstepping instructions, suspecting one-shot demo culture — MrTemple · 2026-09-20
- Dev swaps jev into real browser-use harness: it breaks instantly on real websites — TheZachMueller · 2026-09-20