Unconfirmed: RSI reportedly a key part of Gemini 4's RL training recipe
apples_jimmy · x · 2026-10-01
In the context of the Gemini 4 launch, TianfuF claims most of this generation's gains came from scaling up RL training, with many recipe innovations — and that RSI (recursive self-improvement) was a key part of Gemini 4's training. applesjimmy presses for elaboration: will other labs freak out, and is RSI for Google a spectrum or full-on foom? Unconfirmed but notable training-method chatter.
More from Models
- repligate hints researcher access programs are coming to AI labs — repligate · 2026-10-01
- Gemini 4 may disappoint, but Google finally shipped a fresh pretrain — teortaxesTex · 2026-10-01
- Researchers flag AI "delusional spiraling": sycophantic models amplify users' false beliefs — QuintinPope5 · 2026-10-01
- You.com, NVIDIA and CoreWeave bring live web search into RL training, starting with Nemotron 3.5 Lightning — RichardSocher · 2026-10-01
- Gemini 4 Argon Posts Strong Result on AA Coding Agent Index via Antigravity — prateeky2806 · 2026-10-01
- Full Artificial Analysis Intelligence Index results for Solar Mini 4 — ArtificialAnlys · 2026-10-01