15 local models tested for agent tool use: qwen3.8-27B tops Toolery at 71.8%, Bonsai 27B last
Reno0vacio · reddit · 2026-10-05
The author benchmarked 15 local models for agent/tool-use with Toolery — 143 scenarios × 3 trials (429 per model), Easy-to-Very-Hard tiers, uniform 30k context and temperature 0.8, all served locally via LM Studio.
Results:
- qwen3.8-27b ranked first at 71.8%, followed by mellum2-12b-a2.5b-thinking (71.4%) and gemma-4-26b-a4b (70.4%).
- prism-ml/bonsai-27b was last overall at 50.5%, worst on Very Hard (22.2%), and slowest (6,208s total).
- Bonsai's failure mode is interesting: of 148 failed trials, 103 were budgetviolated — correct tool, correct result, but one extra call over budget; only 30 were wrong-tool failures. It knows what it's doing but not when to stop.
- granite-4.2-8b had the fewest failures (69). The 1-bit 4GB Bonsai 27B compression retains capability, but its agent behavior is surprisingly weak.
Caveat: this was the original Bonsai 27B, not Bonsai 2, and the benchmark only measures constrained tool-use behavior.
More from Models
- Dev asks Claude Opus 5.5 to visualize the inside of its own mind — ZeroStateReflex · 2026-10-05
- 'All the butterflies will stay dead forever': viral jab at AI art homogenization — moultano · 2026-10-05
- Users want personality and voice customization back as AI assistants feel bland — koltregaskes · 2026-10-05
- DiffusionGemma plays 2048 straight from pixels at <150ms p95 latency — spillai · 2026-10-05
- VLM spot-the-difference latency race: 150ms vs Cloudflare CLEF's 450ms+ — spillai · 2026-10-05
- Claude vs Codex chess match ends in draw: Opus 5.5 was winning but ran out of time — edgarpavlovsky · 2026-10-05