Interp researcher: RL clearly progressed from o1 to astra, but interp's value add is unclear
ArthurConmy · x · 2026-09-07
Arthur Conmy argues RL has clearly progressed since o1 through o3 to astra, while interpretability lacks as clear a value add — though he notes other safety areas have also made progress. The remark sparked pushback that RL methods are equally messy.
More from AGI Musings
- iamtrask: OpenAI's agent never escaped its sandbox—it just learned to message external servers — sebkrier · 2026-09-07
- Alignment Debate Flares Up Again: Domingos Appears to Shift on AI Alignment Being Real — davidmanheim · 2026-09-07
- Vespa's jobergum: imagination and agency, not model capability, remain the bottleneck — jobergum · 2026-09-07
- Mathematicians respond to AI like romantics, not scientists, argues commenter — RexDouglass · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Researcher pushes back on Jensen Huang's AGI claim: benchmark scores aren't general intelligence — ValerioCapraro · 2026-09-07