Benchmark: Qwen3.8-27B leads small model agent capabilities
karminski3 · x · 2026-08-31
Benchmark results for small model agent capabilities. The Qwen3.8-27B-UD-Q4KXL version performs best on a single H100 with llama.cpp. Recommended config: reasoningeffort=low for agent tasks, medium/high for coding. Mac users should prefer MLX, as MTP is faster than llama.cpp on macOS.
More from coding & agent
- DIY visual diff tool using GitHub Artifacts and pure JavaScript to save costs — zeeg · 2026-09-01
- Dev shares a dirt-cheap approach to visual diffs — zeeg · 2026-09-01
- Grok Bot automates Shopify updates and supplier coordination — billyjhowell · 2026-09-01
- Grok Bot automates lost deal analysis by mining call and email threads — lennysan · 2026-09-01
- Design pattern: immutable agent artifact revisions behind a stable review URL — RocketSeven · 2026-09-01
- Building a long-term memory benchmark for agents: what to add? — True_Mongoose_7073 · 2026-09-01