Researcher: Frontier Models May Already Have RL-Trained on Auto-Research Tasks, Self-Improvement Here

generativist · x · 2026-09-13

AI researcher generativist re-ran his old auto-research ML agent benchmark with the new astra model and was stunned by the results. He previously argued that when building auto-research-style ML agents with frontier models, it's hard to avoid the conclusion that they've already done remarkable RL on exactly this task — self-improvement is probably already here.

He adds that his harness selects for off-beaten-path solutions, and astra's outputs were "wild," further reinforcing that judgment.

Related event: Researcher Reruns Auto-Research Benchmark, Finds Stunning Results(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →