Benchmark: Tool Calling Performance of Qwen 35B-A3B Variants

OsmanthusBloom · reddit · 2026-08-26

The author benchmarked the tool-calling capabilities of Qwen3.6-35B-A3B and its fine-tunes. Using the tool-eval-bench 2.6.0 suite (88 tests in Hardmode) on a cluster of 32GB V100s, the results showed that Ornith 1.5 and Tiel-Coder (based on Ornith) were the top performers, scoring close to Qwen3.8-27B and significantly higher than Qwen3.6-27B. KAT Coder also slightly outperformed the original 35B-A3B. The tests utilized llama.cpp with Q4 quants across multiple runs, accounting for context pressure.

Original post →

More from Models

Models channel →