Community quants for Qwen3.8 Flash save 20-30GB at same quality as unsloth
Dutchnamn · reddit · 2026-08-29
A Reddit user released a set of Qwen3.8 Flash (Next) GGUF quants after days of benchmarking. They require 20-30GB less disk and RAM than comparable unsloth or AesSedai quants at equal quality. PPL is published in the readme and is competitive; both Q4 quants are strong, with Q3/Q5 to follow. Built with imatrix and a tailored quantization recipe. A ROCmFP4 variant for AMD users is slightly better and faster than Q4XS.
More from Infra
- Scobleizer: Qwen Cloud is natively built around AI for better agent integration — Scobleizer · 2026-08-29
- NVL72 Achieves Up to 30x Better Throughput per MW than GB300 on AgentX Benchmark — nvidia · 2026-08-29
- a16z Partner: Only 2% of US Electricians Certified for DC Power — GregCook2011 · 2026-08-29
- The 'Boring' Network That Saves GPU Training Runs: OOB Management Explained — AccBalanced · 2026-08-29
- Running Generalist Robot Policies on STM32 and ESP32 Chips — yacineMTB · 2026-08-29
- Privacy architecture builds user trust to share sensitive health data with AI — bgmshana · 2026-08-29