Qwen3.8-27B Quantization Analysis: FP8 vs 4-bit Loss

pbaylies · x · 2026-08-20

Spent 67 hours of model time benchmarking Qwen3.8-27B quantization on 4x RTX 3090s. Compared FP8, NVFP4, AWQ INT4, GGUF Q4KM, and NInfer across 4,800 tasks, 10,120 requests, and 14.5M reasoning tokens. The results were surprising.

Original post →

More from Models

Models channel →