Developer finds Luna 6 is a 'massive downgrade' from Luna 5.6 despite better benchmarks
skilliard7 · reddit · 2026-09-23
A developer maintaining an agentic reporting tool reports that Luna 6 is a "massive downgrade" over Luna 5.6, despite slightly better benchmarks and a 50% lower price. In A/B tests at high effort level:
- On every eval involving reviewing written reports and finding relevant items, Luna 6 missed one or more relevant items while Luna 5.6 did not;
- Luna 6 answers in as few tokens as possible, e.g. replying "I found 7 items" without listing them, forcing follow-up questions;
- Across many use cases, the author could not find a single instance where Luna 6 beat 5.6.
He is sticking with Luna 5.6. The takeaway: benchmark scores can diverge badly from real agentic workflow performance — always A/B test on your own tasks before upgrading.
More from coding & agent
- What would an ideal plugin editing environment inside ChatGPT/Codex look like? — Former_Worldliness70 · 2026-09-23
- boxd_sh praised as a major unlock for AI agents needing persistent computers — tobowers · 2026-09-23
- Open-source AI intraday trading bot ranks Nifty 50 every 15s and trades via Zerodha Kite — IndraVahan · 2026-09-23
- Debugging agent planners: which state do you actually save to reproduce a bad plan? — fishyguy3123 · 2026-09-23
- Yoav Goldberg: some tasks just need deterministic rules — agents can write them — yoavgo · 2026-09-23
- Devs migrate from GPT-6 Astra to Opus 5.5 as coding model race churns — rudrank · 2026-09-23