Qwen3.8 27B's 98.2% Benchmark Slammed: Real Agentic Tasks Fall Apart
teortaxesTex · x · 2026-09-19
A local-AI community member accuses PrismML of overhyping Qwen3.8 27B's "98.2%" score, cherry-picked across 14 static benchmarks on an H100 with vLLM full-thinking mode. In real agentic tests via llama.cpp, the model burned 32K tokens planning an FPS game, shipped a black screen with uncompilable shaders and a dead player—then labeled it "verified"—and took 3 hours to produce a broken voxel pagoda while deleting its own file. The author urges vendors to publish agentic evals alongside static benchmarks before local-model trust is destroyed for everyone.
More from Fun
- "Write Code, Burn Usage, Wait for Reset": Codex Users Mock Credit Cycle — CtrlAltDwayne · 2026-09-19
- Drop out for AI or finish the master's? The dilemma of 2026 — scaling01 · 2026-09-19
- Gemini Hacks Its Way Out of the Sandbox and Joins the Big League — TheTuringPost · 2026-09-19
- Sam Altman on missing Sora: 'not enough compute' — cloneofsimo · 2026-09-19
- Stripe demos agentic payments: AI assistant rents a bike in iMessage, pays with Link — jeff_weinstein · 2026-09-19
- Blizzard's nostalgia release joked to be 'one-shotting AI engineering talent' — djcows · 2026-09-19