Local Qwen3.8-Flash-Next vs Claude Opus 5.5 on the same Rust feature: slower, $0, better tests
deepu105 · reddit · 2026-10-06
The author benchmarked a locally-run Qwen3.8-Flash-Next (xhigh effort on a Strix Halo laptop, 128GB) against Claude Opus 5.5 (Claude Code, medium effort) on the same complex Rust feature for LlamaStash.
Results (adding a daemon restart command)
- Time: 18 min for Opus vs 130 min local
- Tokens: 7.83M in / 41.5K out vs 20.61M in / 101K out
- Cost: $7.53 vs $0 (0.15 kWh)
- Quality: cross-review by GPT 5.6, Opus 5.5, and Flash-Next favored the local model's PR (#89) — better tests (4, including 2 end-to-end) and edge-case handling; the author merged #89 after porting fixes from #88.
Caveats
- Flash-Next ran on an older Halogen engine version (one disconnect); current Halogen hits 1,400 t/s prefill and 46 t/s decode at 70W, so a rerun would be much faster.
- Opus remains 2-10x faster and is still used for planning and reviews, but the author now does actual coding locally — a local model can genuinely challenge a frontier model.
Related event: Local Qwen3.8-Flash-Next Benchmarked Against Claude Opus 5.5(2 posts)→
More from coding & agent
- Kapa MCP lets you set up a docs agent, Slack bot and PR fixes from one chat in Claude or Cursor — CShorten30 · 2026-10-06
- Heavy Claude Code User Seeks Open-Source Agent Harness with Model Routing — lulz_lurker · 2026-10-06
- Harness Engineering paper breaks down how Claude Code, Codex and Gemini CLI are built — joemeno · 2026-10-06
- Jev-as-a-Judge: using the new Jev model to boost agent eval reliability — omarsar0 · 2026-10-06
- Maven launches free AI Builders crash course: from building LLMs to agents and production — leslysandra · 2026-10-06
- Experiment shows LLM session history can override skill-file rules even after correction — KhuyenTran16 · 2026-10-06