Unconfirmed: RSI reportedly a key part of Gemini 4's RL training recipe

apples_jimmy · x · 2026-10-01

In the context of the Gemini 4 launch, TianfuF claims most of this generation's gains came from scaling up RL training, with many recipe innovations — and that RSI (recursive self-improvement) was a key part of Gemini 4's training. applesjimmy presses for elaboration: will other labs freak out, and is RSI for Google a spectrum or full-on foom? Unconfirmed but notable training-method chatter.

Original post →

More from Models

Models channel →