Rerunning autoresearch 6 months later: 5 new wins, models are ~10x smarter

MParakhin · x · 2026-09-21

Dev MParakhin reran his autoresearch experiment with identical setup, using GPT-6 Astra and Fable 5.1 at Max effort. Six months earlier, GPT-5.4 xhigh produced only 1 improvement across 103 experiments — still a "free" win for someone with a day job. This time the models delivered 5 additional improvements on top, each exponentially harder to find, leading him to estimate models are now 10x smarter than half a year ago.

Original post →

More from Models

Models channel →