Dream-RSI: agents self-improve policies by 'dreaming' over replayed exploration histories

inductionheads · x · 2026-09-16

A study shared by Steve Hsu proposes Dream-RSI, a recursive self-improvement loop: the agent continuously collects discovery histories via online exploration, builds replay simulators from them, refines meta-exploration strategies through 'dreaming,' then redeploys the improved policy online. Experiments show gains in both discovery effectiveness and efficiency across several settings.

Related event: Google Open-Sources Dream-RSI: Recursive Self-Improvement via Dream-Style Replay(14 posts)→

Original post →

More from AGI Musings

AGI Musings channel →