GPT-6 sees only modest gains on empirical economics: researchers report the 'march of nines'
soumitrashukla9 · x · 2026-09-09
Economists are sharing first impressions of GPT-6. Chris Blattman calls it a step change but a tiny one — roughly the gap between 5.5 Codex and 5.5 ChatGPT Pro — occasionally worth using, but no dramatic improvement.
Yanagizawa reports similar results at Project APE: a key benchmark moved from 99% to 99.5%, the start of a "march of nines" where each additional nine signals a steep drop in error rate despite little visible jump.
Related event: Economists Find GPT-6 Upgrade Marginal and Barely Noticeable(2 posts)→
More from Models
- GPT-6 Astra tops RSI-Exam at 0.5126, 18.4% above GPT-5.6 Sol — HuaxiuYaoML · 2026-09-09
- APEX-Agents 1.1 benchmark update: Claude Fable 5.1 tops leaderboard at 68.6% — amaarora · 2026-09-09
- Frontier labs must shrinkflate the $200 subscription to upsell you to API pricing — StewartalsopIII · 2026-09-09
- User math: peak pricing 2x but base rate better, 252M cache tokens cost just $0.75/day — teortaxesTex · 2026-09-09
- Astra noticeably worse than Sol in long threads, dev finds handoff workaround — jdjohnson · 2026-09-09
- VC communism is over: frontier models on rationing force hard model-choice thinking — StewartalsopIII · 2026-09-09