Two GGUF Quantization Repositories and Benchmarks

Blahblahblakha · reddit · 2026-07-11

Two GGUF quantization repositories are shared, along with reproducible quantization and performance data: - **Hy3 (Tencent 295B MoE, 21B active)**: Offers multiple quantization versions like Q6_K, Q4_K_M, Q3_K_L, and IQ2_M, detailing their respective sizes, Mean KLD, top-token consistency, and generation speed. The author recommends **Q4_K_M** as a more practical default; with around 256GB of VRAM/RAM, **Q6_K** can be used with virtually no loss. - **Nemotron-Labs-Audex-30B-A3B (NVIDIA)**: Provides both text and audio quantization tracks, including text-only GGUF, audio_quants, and audio sidecar components. The text provides quality metrics and throughput data across three corpora, showing that quantized speeds are roughly twice as fast as BF16. Additionally, a few usage caveats are highlighted: - For Hy3 on CUDA decode, `--split-mode layer` is required; tensor split will crash. - Nemotron-Labs-Audex-30B-A3B inherits the **NVIDIA OneWay Noncommercial** license and cannot be used commercially. - The complete audio pipeline currently still requires the sidecar and NVIDIA scripts; there is no "all-in-one" audio GGUF runtime yet.

Original post →

More from Infra

Infra channel →