Local 27B Face-off: Dirk-Qwen3.8 Beats Swift-1.5 on a 200-Question Personal Eval
norenEnmotalen · reddit · 2026-10-01
The author ran his own domain-specific eval set (200+ questions covering coding, numpy/pandas, data analytics decisions, local RAG and voice assistant tasks) on an M1 Max 32GB to compare two quantized models built on the same Qwen3.8-27B base.
- Setup: Dirk-Qwen3.8-27B-UD-Q4KXL (128K context) vs Swift-1.5-Qwen3.8-27B-Q4KL (110K), on a modified Splash inference stack
- Surprising result: Dirk was decisively sharper — faster, fewer tokens, more correct answers. Swift-1.5 often churned on hard questions until hitting the 16,384 max-token truncation limit
- Even excluding truncation failures, Dirk wins on token efficiency; the author ruled out prompt-caching artifacts
- A key factor: the "be brief" instruction injected via the chat template proved more impactful than expected, and the pattern holds on real code refactoring tasks
- Aside: the only model to pass 100% of his eval packs is Claude Opus 5.5; Deepseek Flash 4.1 fp32 missed just three questions
More from Infra
- Polymarket prices 26% odds a US state enacts a data center moratorium by 2026 — Polymarket · 2026-10-01
- Google's Spanner Omni goes GA with 2M+ downloads, bringing distributed SQL to any cloud or laptop — rakyll · 2026-10-01
- Meta Claimed $3.9B Research Tax Credit by Labeling AI Data Centers as 'Experiments' — mkheck · 2026-10-01
- Ben Lorica: the AI data problem didn't disappear — it moved downstream into permissions and pipelines — bigdata · 2026-10-01
- Memory's $200B inflection: concurrent AI sessions turn DRAM into an architecture problem — BenBajarin · 2026-10-01
- Google VP: a fraction of campus philanthropy could fund compute cloud for all university researchers — jasondeanlee · 2026-10-01