Dream-RSI: agents self-improve policies by 'dreaming' over replayed exploration histories
inductionheads · x · 2026-09-16
A study shared by Steve Hsu proposes Dream-RSI, a recursive self-improvement loop: the agent continuously collects discovery histories via online exploration, builds replay simulators from them, refines meta-exploration strategies through 'dreaming,' then redeploys the improved policy online. Experiments show gains in both discovery effectiveness and efficiency across several settings.
More from AGI Musings
- Researcher: AI math 'darlings' long relied on fake baselines, math lacks empirical tradition — RexDouglass · 2026-09-17
- AI slop isn't bad writing—it's unchecked content; a 10-second sniff test beats detectors — thisdudelikesAI · 2026-09-17
- Gowers responds to letter on maths and AI signed by 25 Fields medallists — tak3sh8 · 2026-09-17
- mark_k calls for a documentary exposing Effective Altruism's grip on AI doomerism — mark_k · 2026-09-17
- The overlooked existential cost of AGI: the loss of meaning in intellectual effort — ferruz_noelia · 2026-09-17
- Ex-Anthropic researcher on CBS: AI could copy itself across machines before you unplug it — elonmusk · 2026-09-17