GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both
mhmazur · x · 2026-09-26
User mhmazur visualized reasoning token usage of GPT-5.6 vs GPT-6 Luna on ARC-AGI-2 at max effort: x-axis for 5.6, y-axis for 6, with each half of a circle showing correctness.
- Most circles fall below the diagonal, confirming GPT-6 used fewer reasoning tokens than 5.6 on the majority of tasks.
- Red dots in the top-right mark the harder, token-heavy tasks — which both models mostly got wrong, exposing a shared ceiling where extra compute doesn't buy correctness.
Note: the model version names are unverified.
More from Models
- Opus 5.5 unblocks a 4-month ts-rust port in 10 hours where GPT-6 Astra stalled at 85% — EricBuess · 2026-09-26
- Kev-4B is now available on OpenRouter — charles_irl · 2026-09-26
- GPT 6 Luna Max one-shots porting an Nvidia project to wgpu shaders — and it's faster — mgostIH · 2026-09-26
- Observation: Astra uses filler tokens far more effectively than other models — scaling01 · 2026-09-26
- How Long Until Local ~30B A3B Models Match GLM 5.3 Flash Quality? — Aggravating-Push-207 · 2026-09-26
- Ethan Mollick: 'Keep prompts short' is bad advice, and minimizing token cost confuses inputs with outputs — emollick · 2026-09-26