Testing 5 Qwen3.6-35B-A3B finetunes: the base model beats almost all of them
returnity · reddit · 2026-09-29
Reddit user returnity benchmarked the Qwen3.6-35B-A3B base model against five finetunes (Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, Nex-N2.5-mini) on Aider Polyglot:
- Base wins nearly across the board: 37.4% first-try pass, 71.0% retry pass, 96.3% well-formed diffs
- Only Occamy-1.0 is competitive (30.8%/69.2%) while using fewer tokens
- KAT-Coder is far less accurate (20.6%) but extremely fast and token-efficient — a good subagent
- Community favorite Tiel-Coder disappoints (18.7%); its 'terse mode' sharply cuts accuracy
- Lowering reasoning effort (xhigh→medium) cuts tokens/time by 1/3 with little accuracy loss
- Nex-N2.5-mini is worst at 10.3% first-try and 33.3K tokens per solve
Chat template choice had minimal effect on scores.
More from Models
- Early hands-on says Sonnet 5.5 looks benchmaxxed: pricier and more token-hungry than Sonnet 5 — haider1 · 2026-09-29
- Japan's sovereign LLM project LLM-jp releases open 4.1 models with tool calling — markjeffrey · 2026-09-29
- Claude Code Projects defaults to low effort, and an Anthropic engineer teases Sonnet 5.5 — lydiahallie · 2026-09-29
- Anthropic Sonnet 5.5 Debuts at No. 2 on Vals Index, Just 0.47 Points Behind Opus 5.5 — airesearch12 · 2026-09-29
- Claude suddenly stopped cheating, user tests show a behavior shift — basedjensen · 2026-09-29
- 'My own personal move 37': users report Claude and Fable 5.1 proposing genuinely good ideas they missed — zetalyrae · 2026-09-29