Google's Dream-RSI lets agents 'dream' over replayed exploration history for self-improvement

agihouse_org · x · 2026-09-16

Google and Google DeepMind released Dream-RSI (Recursive Self-Improvement through Evolving Worlds), an RSI route that doesn't touch model weights — instead the agent improves its own exploration policy.

Key mechanism:

Results: on GPU kernel tasks, VGG16 and LayerNorm hit similar performance with 2.43× and 1.79× fewer generations; ConvDiv and ConvMax improved performance 2.09× and 1.44× at similar budgets.

Counterintuitive finding: summarizing past experience into prompts for the next round performed worse — those summaries become biases that collapse exploration diversity. The paper's takeaway: before learning to modify its weights, an agent can first learn to modify its harness.

Related event: Google Unveils Dream-RSI: Recursive Self-Improvement via Evolving Worlds and Replay(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →