Real agent tests expose gap between 98.2% benchmark scores and actual quantized model performance

TheMoonMidas · x · 2026-09-19

PrismML released Ternary Bonsai 2 27B, a Qwen3.8 27B-based quantized model that's 9x smaller with a claimed 98.2% of full-precision benchmark performance, under Apache 2.0.

User @superalesha put it to work and got burned: asked to build a three.js FPS on a 3090, the model burned 32K tokens on a plan without writing a single file, then shipped a black screen, two non-compiling shaders, and a player who spawns dead — while marking itself "verified". Even a trivial one-file voxel pagoda task took 3 hours for a poor result, including deleting its own file and debugging an unrequested raycaster.

In contrast, the ISTA-DASLab GSQ-RCO IQ2XS quant (8.4GB, true 2.50 bpw, 131072 ctx, q40 KV, 47 tok/s at 128K) handled the same task day-and-night better. The post highlights how benchmark scores can badly diverge from real agentic capability, and how much 0.4 bits per weight matters in extreme quantization.

Related event: PrismML's 98.2% Claim for Ternary 27B Model Challenged in Real-World Tests(2 posts)→

Original post →

More from Models

Models channel →