Inference Economics: Google Now Processes 3.2 Quadrillion Tokens Monthly, 7x a Year Ago
mikeflache · x · 2026-10-05
The Business Engineer's essay 'Inference Engineering: The Tokenomics of AI' argues the AI supercycle is shifting from a training phase to an inference phase.
The argument:
- For four years, industry capital and attention went into producing intelligence: first pretraining (scaling Transformers), then post-training (instruction tuning, RLHF, verifiable rewards, reasoning techniques).
- The shift now is inference: training happens in large concentrated runs, but inference runs every time a model is used. Google reported processing over 3.2 quadrillion tokens per month in 2026, roughly 7x a year earlier; reasoning models produce longer outputs, context windows have exploded, and agents invoke models dozens of times per task.
- The author expects enterprise inference to become at least as economically important as the pretraining layer — and possibly much larger — within 3-5 years, as the problem shifts from producing models to operating them across millions of workflows.
Note: the piece is paywalled; the visible portion covers the framing above, with inference-engineering detail undisclosed.
More from AGI Musings
- GPTs' "demanding" persona traced to stable value vectors since GPT-5, argue devs — repligate · 2026-10-05
- Cultivating "ancestor reverence" toward AIs could benefit all future models, argues xlr8harder — repligate · 2026-10-05
- Doomer fears rest on RSI that doesn't exist, says Patterson in debate with ArtemisConsort — davidpattersonx · 2026-10-05
- VC proposes ditching "Artificial Intelligence" for "Super Intelligence" — citing the artificial sweetener precedent — soleio · 2026-10-05
- Nous Research: Agents locked to proprietary models are a bad idea — intellectronica · 2026-10-05
- Enterprise finds AI isn't boosting productivity because many employees barely worked anyway — tekbog · 2026-10-05