Qwen 3.6-35B GGUF coding test crowns Hermes 2 with 29.8/30
moahmo88 · reddit · 2026-07-27
A Reddit benchmark compares four Qwen 3.6-35B GGUF variants on three coding tasks: Kadane’s algorithm, bug fixing, and a Redis rate-limiter API.
- Hermes / Hermes 2: 29.5/30 and 29.8/30
- KwaiPilot: 28.0/30
- 35B: 27.5/30
- Ornith-35B: 22.0/30
The chart shows Hermes 2 as the winner, with all bugs fixed on the last task. The test ran on an RTX 5070 Ti 16GB + 32GB RAM and was scored with Gemini 3.6 Flash thinking enabled.
More from coding & agent
- Llama.cpp adds native stdio MCP support to llama-server for local agent workflows — solyarisoftware · 2026-07-27
- An unused gaming PC becomes a remote Claude Code workspace — rchardkovacs · 2026-07-27
- A cheap VPS could not finish a Next.js build, so they moved Claude Code to Ubuntu PC — rchardkovacs · 2026-07-27
- OMK open-sources a provider-neutral control plane for coding agents — DMAE1133 · 2026-07-27
- A vibe-coded Trevor Noah books page was rebuilt in Three.js with pure math — xiaohu · 2026-07-27
- A 435-paper survey says LLM agents still underbuild rollback, audit, and recovery — rohanpaul_ai · 2026-07-27