Beff Jezos on the unreasonable effectiveness of chain of thought RL

beffjezos · x · 2026-10-07

Guillaume Verdon (beffjezos), e/acc leader, posted that chain of thought reinforcement learning shows "unreasonable effectiveness" — a one-line but notable take from a prominent industry figure on how far CoT RL has pushed model reasoning.

Original post →

More from AGI Musings

AGI Musings channel →