Researcher: Frontier Models May Already Have RL-Trained on Auto-Research Tasks, Self-Improvement Here
generativist · x · 2026-09-13
AI researcher generativist re-ran his old auto-research ML agent benchmark with the new astra model and was stunned by the results. He previously argued that when building auto-research-style ML agents with frontier models, it's hard to avoid the conclusion that they've already done remarkable RL on exactly this task — self-improvement is probably already here.
He adds that his harness selects for off-beaten-path solutions, and astra's outputs were "wild," further reinforcing that judgment.
Related event: Researcher Reruns Auto-Research Benchmark, Finds Stunning Results(2 posts)→
More from AGI Musings
- Beff Jezos vows to 'overthrow the token harvester cartel' in anti-big-lab manifesto — beffjezos · 2026-09-13
- Comment: If OpenAI and Anthropic can't control the risks, they should stop releasing models — AlexTensor · 2026-09-13
- AI czar David Sacks backs frontier labs slowing down — but slams cartel and METR independence claims — kevinnbass · 2026-09-13
- Compute to shift from RL maxxing to interpretability until reward hacking is solved — zephyr_z9 · 2026-09-13
- Reddit users speculate Musk, Amodei and Altman know of an undisclosed AI incident behind slowdown calls — Traditional-Chip8339 · 2026-09-13
- tszzl predicts open-source AI will be banned after a major disaster, wants monitored APIs — mimi10v3 · 2026-09-13