GPT-6 sees only modest gains on empirical economics: researchers report the 'march of nines'

soumitrashukla9 · x · 2026-09-09

Economists are sharing first impressions of GPT-6. Chris Blattman calls it a step change but a tiny one — roughly the gap between 5.5 Codex and 5.5 ChatGPT Pro — occasionally worth using, but no dramatic improvement.

Yanagizawa reports similar results at Project APE: a key benchmark moved from 99% to 99.5%, the start of a "march of nines" where each additional nine signals a steep drop in error rate despite little visible jump.

Related event: Economists Find GPT-6 Upgrade Marginal and Barely Noticeable(2 posts)→

Original post →

More from Models

Models channel →