GPT-5.6 Still Generates Too Many Reasoning Tokens

bdsqlsz · x · 2026-07-10

The author's comparison tests suggest that Luna is roughly equivalent to GPT-5.5 mini, Terra to GPT-5.5 medium, and Sol performs the best—with even lower-tier Sol outperforming the highest-tier Terra.

The quote adds that while GPT-5.6 sol is "smarter," it still suffers from excessive reasoning token generation. Unmodified, it consumes 518n-2 reasoning tokens even when the answer is correct. After the author's modifications, it used fewer tokens and achieved a 5/5 accuracy rate.

Related event: GPT-5.6 Release Sparks Discussion on Performance and Cost(10 posts)→

Original post →

More from Models

Models channel →