Interp researcher: RL clearly progressed from o1 to astra, but interp's value add is unclear

ArthurConmy · x · 2026-09-07

Arthur Conmy argues RL has clearly progressed since o1 through o3 to astra, while interpretability lacks as clear a value add — though he notes other safety areas have also made progress. The remark sparked pushback that RL methods are equally messy.

Related event: Researchers Debate the Value and Boundaries of Mechanistic Interpretability(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →