Local Qwen3.8-Flash-Next vs Claude Opus 5.5 on the same Rust feature: 7x slower, free, better tests
deepu105 · reddit · 2026-10-06
A developer ran the same vague-prompt Rust task — adding a llamastash daemon restart command reusing existing start/stop code — on local Qwen3.8-Flash-Next (xhigh, Pi harness, ASUS ROG Flow Z13 / Strix Halo 128GB) and cloud Claude Opus 5.5 (medium, Claude Code):
- Time: 18 min vs 130 min; tokens: 7.83M/41.5K vs 20.61M/101K; cost: $7.53 vs 0.15 kWh.
- Quality: Flash-Next added 4 tests (2 end-to-end) and handled edge cases better. Cross-review by GPT 5.6, Opus 5.5, and Flash-Next: two picked the Flash-Next PR; the author merged #89 after porting fixes from #88.
- Caveats: Opus ran at medium effort; Flash-Next used an older Halogen version that dropped once — the current version does 1,400 t/s prefill and 46 t/s decode at 70W, so a redo would be much faster.
Verdict: Opus is still 2–10x faster and remains the choice for planning and reviews, but actual coding now happens locally — a laptop model can genuinely challenge a frontier model.
Related event: Local Qwen3.8-Flash-Next Benchmarked Against Claude Opus 5.5(2 posts)→
More from coding & agent
- Claude Opus 5.5 took 90,000 screenshots of its own game over two days to polish the visuals — prasenx · 2026-10-06
- classif: shell scripts branch on meaning via one token's logprobs on a local 12B — piotr1215 · 2026-10-06
- Agent 3D-renders its own game assets to build Star Control Hyper Melee unsupervised — draginol · 2026-10-06
- Markov: a pi-style, transparent LLM harness in a single Bash script with sub-agents — biller23 · 2026-10-06
- Meta paper: dual coding agents cross-review lift correct patches from 45.8% to 62.5% — rohanpaul_ai · 2026-10-06
- Rewriting from scratch beats depending on a single maintainer with feelings — l4rz · 2026-10-06