Local agent eval: challenger model wrote tool calls as prose 5/80 times — a silent failure class

Grimmoner · reddit · 2026-09-27

A developer ran a proper comparison before swapping models on local agent seats: 20 real agent tasks (read/write/edit/patch), each run across 4 seed/thinking modes for 80 attempts per model, pitting qwen3.5:9b against MiMo-V2.6-Distill-Qwen-9B at IQ4XS. Score: 45/80 vs 30/80 strict passes, Fisher exact p=0.026 — a real gap.

But the failure shapes mattered more than the score:

Methodology note: the author byte-compared both GGUFs' chat templates first, ruling out the most common cause of "bad tool calling" claims. Caveats: single turn, one quant, no reasoning eval. Takeaways: count text-form tool calls as a distinct failure class, and assert on the tool list.

Original post →

More from coding & agent

coding & agent channel →