Benchmark: Tool Calling Performance of Qwen 35B-A3B Variants
OsmanthusBloom · reddit · 2026-08-26
The author benchmarked the tool-calling capabilities of Qwen3.6-35B-A3B and its fine-tunes. Using the tool-eval-bench 2.6.0 suite (88 tests in Hardmode) on a cluster of 32GB V100s, the results showed that Ornith 1.5 and Tiel-Coder (based on Ornith) were the top performers, scoring close to Qwen3.8-27B and significantly higher than Qwen3.6-27B. KAT Coder also slightly outperformed the original 35B-A3B. The tests utilized llama.cpp with Q4 quants across multiple runs, accounting for context pressure.
More from Models
- Together Ranks Top Open Models: Kimi K3 and DeepSeek V4 Lead Use Cases — togethercompute · 2026-08-26
- Questions over Astra's progress: 2 months for 3 more models? — teortaxesTex · 2026-08-26
- View: Tokens-per-second matters more than model size now — natesiggard · 2026-08-26
- Tiel-Coder-35B achieves 121.4 tok/s for local inference — DerTomsn · 2026-08-26
- Grok 4.6 now available on OpenCode Go subscription — veggie_eric · 2026-08-26
- One Claude Design Task Burned 95% of the $100/Mo Plan — zeeg · 2026-08-26