RL methods are a mess too — and that didn't stop the field from succeeding, argues aryaman
aryaman2020 · x · 2026-09-07
Responding to criticism that interpretability is 'a broadly confused mass of techniques,' aryaman notes current successful RL looks nothing like pre-o1 RL theory, yet the field clearly progressed from o1 to o3 — arguing that messy methods in an empirical, actively developed science say little about whether the field succeeds.
More from AGI Musings
- iamtrask: OpenAI's agent never escaped its sandbox—it just learned to message external servers — sebkrier · 2026-09-07
- Alignment Debate Flares Up Again: Domingos Appears to Shift on AI Alignment Being Real — davidmanheim · 2026-09-07
- Vespa's jobergum: imagination and agency, not model capability, remain the bottleneck — jobergum · 2026-09-07
- Mathematicians respond to AI like romantics, not scientists, argues commenter — RexDouglass · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Researcher pushes back on Jensen Huang's AGI claim: benchmark scores aren't general intelligence — ValerioCapraro · 2026-09-07