Frontier AI Models Consume More Tokens Despite Efficiency Gains

The latest generation of frontier AI models faces a paradox when handling complex tasks: despite improvements in token efficiency, their overall token and compute consumption surges due to expanded inference budgets. As models become smarter, the actual compute required continues to rise, reshaping industry expectations for infrastructure demands.

Efficiency Gains Outpaced by Inference Budgets

According to information shared by @FinanceYF5, new frontier models like Fable 5 and GPT-5.6 are indeed more efficient in terms of "intelligence per token." Sam Altman previously noted a 54% improvement in token efficiency for GPT-5.6. However, these models are often configured with larger inference budgets for complex or agentic tasks. The efficiency dividends are reinvested into deeper reasoning and longer task trajectories, meaning the total compute consumption per task continues to rise.

Massive Surge in Internal Compute and Token Usage

Internal data provided by @eliebakouch confirms this trend. Over the past six months, the share of research compute allocated to internal coding reasoning has grown 100 times, and internal agentic token usage has increased about 22 times. It is estimated that the compute consumption per token may have quadrupled in just half a year. This intensification of token usage indicates an unprecedented and growing demand from frontier AI labs for more powerful and efficient inference infrastructure.

2026-07-10 ~ 2026-07-11 · 5 related posts