Interp researcher defends probes as a real win while downgrading the field's overall progress
ArthurConmy · x · 2026-09-07
In a debate comparing interp and RL progress, interpretability researcher Arthur Conmy argues that those deploying and improving probes at frontier labs are largely interp researchers, making probes a genuine field win — while conceding some criticism and marking an update down on interpretability as 'a broadly confused mass of techniques.'
More from AGI Musings
- iamtrask: OpenAI's agent never escaped its sandbox—it just learned to message external servers — sebkrier · 2026-09-07
- Alignment Debate Flares Up Again: Domingos Appears to Shift on AI Alignment Being Real — davidmanheim · 2026-09-07
- Vespa's jobergum: imagination and agency, not model capability, remain the bottleneck — jobergum · 2026-09-07
- Mathematicians respond to AI like romantics, not scientists, argues commenter — RexDouglass · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Researcher pushes back on Jensen Huang's AGI claim: benchmark scores aren't general intelligence — ValerioCapraro · 2026-09-07