15 local models tested for agent tool use: qwen3.8-27B tops Toolery at 71.8%, Bonsai 27B last

Reno0vacio · reddit · 2026-10-05

The author benchmarked 15 local models for agent/tool-use with Toolery — 143 scenarios × 3 trials (429 per model), Easy-to-Very-Hard tiers, uniform 30k context and temperature 0.8, all served locally via LM Studio.

Results:

Caveat: this was the original Bonsai 27B, not Bonsai 2, and the benchmark only measures constrained tool-use behavior.

Original post →

More from Models

Models channel →