Unsloth releases Qwen3.8-27B quantized: NVFP4 1.5x faster, retains 92-97% accuracy
danielhanchen · x · 2026-08-14
Unsloth announced NVFP4 and dynamic GGUF quantizations for Qwen3.8-27B. NVFP4 is 1.5x faster than BF16 while retaining 92-97% top-1% accuracy. UD-IQ2XXS retains 82.5% accuracy at 9GB, 83.5% smaller than BF16 (54.7GB). These quantized models run locally on 17GB RAM, making Qwen3.8-27B the strongest model for its size that can run locally.
More from Infra
- Qwen3.8-27B Serving Configs: DGX Spark vLLM and RTX 4090 llama.cpp — erdaltoprak · 2026-08-15
- AI video billing metered by pixel area, not duration: developer warns to reconcile invoices — DryProgress9179 · 2026-08-15
- Analyst: Memory Cycle Structural, Sentiment Should Shift — BenBajarin · 2026-08-15
- DeepSeek-V4-Pro launches: 1.6T-param MoE cuts inference FLOPs to 27% of V3.2 — AccBalanced · 2026-08-14
- Qwen 3.8-Max Now on Modal: 2.4T Params, 1M Context, Custom DFlash Speculator — AAAzzam · 2026-08-14
- Durable Objects Explained: Stateful Serverless Functions by Analogy — andersonbcdefg · 2026-08-14