Beff Jezos on the unreasonable effectiveness of chain of thought RL
beffjezos · x · 2026-10-07
Guillaume Verdon (beffjezos), e/acc leader, posted that chain of thought reinforcement learning shows "unreasonable effectiveness" — a one-line but notable take from a prominent industry figure on how far CoT RL has pushed model reasoning.
More from AGI Musings
- AI has now cracked at least 10 open math problems each worthy of a Fields Medal — luismbat · 2026-10-07
- Altman: world should accept some bad things happening for AI's benefits and agency — Duckducklaugh · 2026-10-07
- OpenAI publishes 372 AI-generated math proofs on GitHub, daring academia to keep up — The Decoder · 2026-10-07
- AI Could Master Transcendental Arguments About Itself—They Still Wouldn't Move Us — birchlse · 2026-10-07
- Verifiable domains will all fall to AI like maths — the open question is the rest — charlieharris01 · 2026-10-07
- Shogi AI rates a classic "good for White" line at +800, upending decades of human opening theory — i_dg23 · 2026-10-07