Open-weight models match Fable 5 on hard LiveCodeBench at a tenth of the cost: $5.76 vs $61.11
vecto_r · reddit · 2026-09-18
- A training-free manager-worker scaffold uses fresh instances of one model to decompose problems and coordinate, letting open-weight models approach or match Claude Fable 5 on hard LiveCodeBench.
- Gains from orchestration: Qwen3.8-FlashNext 84.2% → 93.0%; Qwen3.8-27B 69.2% → 92.4%; GPT-5.6-Terra 80.8% → 88.0%; GPT-5.6-Luna 70.4% → 81.2%; vs Fable 5's 90.4% single-call baseline.
- Cost gap is stark: orchestrated FlashNext runs $5.76 per pass against Fable 5's $61.11.
- Code and paper are public (GitHub GVS5H, arXiv 2608.26480).
More from coding & agent
- Teknium: Jev can't compact context well — Hermes summarizes 95% of it away — Teknium · 2026-09-20
- HarnessRouter open-sources a unified API to run Codex, Claude Code and more as agent backends — daniel_mac8 · 2026-09-20
- HarnessRouter: routing agent harnesses instead of models, a fresh infra idea — daniel_mac8 · 2026-09-20
- GameToMac launched 10 days ago and already runs AoE IV, CS2 and Diablo IV on Apple Silicon — nickbaumann_ · 2026-09-20
- DialKit 2.0 ships: open-source real-time UI tuning tool with prompts for coding agents — LinusEkenstam · 2026-09-20
- Jev loses to Gemini on 1,565-email classification benchmark, but dev still wants it in production — socialwithaayan · 2026-09-20