Qwen3.8-27B scores 29/30 on AIME 2026 with FP8, matching frontier flagships

No_Run8812 · reddit · 2026-08-21

A community benchmark of Qwen3.8-27B on MathArena/aime2026 comparing BF16 vs FP8 weights at medium and xhigh reasoning effort:

Against frontier models (single pass@1): GPT-5.6 Sol xhigh reported 99.9%, GLM-5.2 and GPT-5.4 99.2%, Gemini 3.1 Pro 98.3%, Claude Opus 4.6 and DeepSeek V4 Pro 96.7% — a 27B FP8-quantized model matching several flagships.

Settings: exact-match scoring, temperature zero, sampling disabled, identical chat template and prompt format.

Original post →

More from Models

Models channel →