Qwen 3.8 27B scores 72.9 on Aider benchmark, tying Gemini 2.5 Pro
Baldur-Norddahl · reddit · 2026-08-24
The author ran the Aider benchmark with Qwen 3.8 27B FP8 + FP8 KV cache at 256K context on vLLM. Score: 72.9 — matching Gemini 2.5 Pro (2025-04-12, 72.9), beating Claude Opus 4 (2025-05-25, 72.0) and DeepSeek R1 (71.4).
He notes it's an older benchmark, but it's remarkable that a MacBook now matches SOTA models from just over a year ago. Real-world performance is even better since the harness improved: with the DeepSeek Harness it would solve most if not all Aider tests, though often using more than 2 turns.
More from Models
- OpenAI's new 5.6 sol update reportedly beats fable as the best model available — i_dg23 · 2026-08-24
- Google Gemma-4-26B Ported to Apple Silicon with Half Memory Footprint — jasonkneen · 2026-08-24
- McByte sets SOTA on SportsMOT using segmentation masks for tracking — huggingface · 2026-08-24
- Opus 3 vs 5: The evolution of model alignment and hedging — repligate · 2026-08-24
- RTX 5000 Pro Runs Qwen3.8 at Unbeatable Value — Valuable-Run2129 · 2026-08-24
- Anthropic has not upgraded Opus models in over 6 months — tenobrus · 2026-08-24