Rerunning autoresearch 6 months later: 5 new wins, models are ~10x smarter
MParakhin · x · 2026-09-21
Dev MParakhin reran his autoresearch experiment with identical setup, using GPT-6 Astra and Fable 5.1 at Max effort. Six months earlier, GPT-5.4 xhigh produced only 1 improvement across 103 experiments — still a "free" win for someone with a day job. This time the models delivered 5 additional improvements on top, each exponentially harder to find, leading him to estimate models are now 10x smarter than half a year ago.
More from Models
- Codex limits draining 5x faster with 18% less usage, dev's tracking claims — StewartalsopIII · 2026-09-21
- Observation: Models Differ Wildly in Tool-Call Intermediate Steps — DanielLockyer · 2026-09-21
- Rumor: Anthropic's Opus 5.5 landing this week, said to beat Astra on price and quality — daniel_mac8 · 2026-09-21
- Burkov calls stealthy LLM-rival project Jev 'BS' over speed and calibration claims — burkov · 2026-09-21
- Moonshot and Tencent Hunyuan both building Flash models to target agent inference costs — TheZachMueller · 2026-09-21
- Dev Launches Made With Jev, a Free Directory Cataloging Demos, Tools and Skills for the New Model — Sea_Supermarket_5891 · 2026-09-21