Real agent tests expose gap between 98.2% benchmark scores and actual quantized model performance
TheMoonMidas · x · 2026-09-19
PrismML released Ternary Bonsai 2 27B, a Qwen3.8 27B-based quantized model that's 9x smaller with a claimed 98.2% of full-precision benchmark performance, under Apache 2.0.
User @superalesha put it to work and got burned: asked to build a three.js FPS on a 3090, the model burned 32K tokens on a plan without writing a single file, then shipped a black screen, two non-compiling shaders, and a player who spawns dead — while marking itself "verified". Even a trivial one-file voxel pagoda task took 3 hours for a poor result, including deleting its own file and debugging an unrequested raycaster.
In contrast, the ISTA-DASLab GSQ-RCO IQ2XS quant (8.4GB, true 2.50 bpw, 131072 ctx, q40 KV, 47 tok/s at 128K) handled the same task day-and-night better. The post highlights how benchmark scores can badly diverge from real agentic capability, and how much 0.4 bits per weight matters in extreme quantization.
Related event: PrismML's 98.2% Claim for Ternary 27B Model Challenged in Real-World Tests(2 posts)→
More from Models
- 16-model calibration test: open-weight models almost never admit uncertainty — AlexKim · 2026-09-19
- Jev answers in 455ms — a gate you can afford to run on everything — AlexKim · 2026-09-19
- I ran 16 models to vet one tool: one task is not a benchmark — AlexKim · 2026-09-19
- Dev tests 16 models to evaluate TypeSafe's Jev — it ranked 10th on accuracy — AlexKim · 2026-09-19
- Mystery Model Jev Launches Claiming 200x Speed and 400x Cost Cuts, Devs Impressed — multiply_matrix · 2026-09-19
- GLM 5.3 Flash leads quality, Qwen 3.8 Flash Next wins speed in open small-model comparison — HankYeomans · 2026-09-19