Qwen3.8-27B scores 29/30 on AIME 2026 with FP8, matching frontier flagships
No_Run8812 · reddit · 2026-08-21
A community benchmark of Qwen3.8-27B on MathArena/aime2026 comparing BF16 vs FP8 weights at medium and xhigh reasoning effort:
- FP8 xhigh hits 29/30 (96.7%), tying BF16 xhigh while decoding jumps from 28 to 76 tk/s with much faster pre-fill;
- FP8 medium scored 26/30 vs BF16 medium's 28/30 — slight quantization cost at medium effort;
- On problem 7, both xhigh runs exhausted the token budget with no final answer (empty, not wrong).
Against frontier models (single pass@1): GPT-5.6 Sol xhigh reported 99.9%, GLM-5.2 and GPT-5.4 99.2%, Gemini 3.1 Pro 98.3%, Claude Opus 4.6 and DeepSeek V4 Pro 96.7% — a 27B FP8-quantized model matching several flagships.
Settings: exact-match scoring, temperature zero, sampling disabled, identical chat template and prompt format.
More from Models
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24