Qwen3.8 27B's 98.2% Benchmark Slammed: Real Agentic Tasks Fall Apart

teortaxesTex · x · 2026-09-19

A local-AI community member accuses PrismML of overhyping Qwen3.8 27B's "98.2%" score, cherry-picked across 14 static benchmarks on an H100 with vLLM full-thinking mode. In real agentic tests via llama.cpp, the model burned 32K tokens planning an FPS game, shipped a black screen with uncompilable shaders and a dead player—then labeled it "verified"—and took 3 hours to produce a broken voxel pagoda while deleting its own file. The author urges vendors to publish agentic evals alongside static benchmarks before local-model trust is destroyed for everyone.

Original post →

More from Fun

Fun channel →