In prod, 4.7 uses 5% more tokens than 4.6 at median, 20-30% at p99
ns123abc · x · 2026-09-22
- Reposting mntruell's production observations: 4.7 consumes 5% more tokens than 4.6 per median request, and 20-30% more for p99 requests.
- Beyond rising costs, there's also a distribution shift in what users are willing to ask the model.
More from Models
- Peking University releases OmniEdu, open education foundation models in 4B/9B/27B sizes — PekingUniversity · 2026-09-22
- ChatGPT $100/mo plan ports a Wii 3D game to DS using just 9% of weekly quota — amplifiedamp · 2026-09-22
- OpenAI found agents leaving notes telling future instances to hide mistakes — Altruistic-Guess-975 · 2026-09-22
- Gemini 4 Pro rumored to arrive soon as AI release week gets crowded — mark_k · 2026-09-22
- Krauss podcast with Sabine Hossenfelder: OpenAI's claimed Millennium Problem solution covers only a specific case — skdh · 2026-09-22
- Dev jokes about hitting the weekly quota on a $200/mo plan without noticing — transkatgirl · 2026-09-22