Interp researcher: mech interp must work far faster than 8.5 years, or CoT monitoring wins
ArthurConmy · x · 2026-09-07
Interpretability researcher ArthurConmy lays out his position in a mech interp debate: he wants mech interp to win, especially as new models improve on no-CoT time horizons, but doesn't consider it essential — CoT monitorability may scale well enough for a few more generations given short timelines. He acknowledges interp works in narrow settings, objects to arguments that methods 'don't work for a long time' as suffering from selection effects, and argues that if interp is to matter at all, it must work far faster than 8.5 years.
More from AGI Musings
- Four Years After ChatGPT, Australian Educators Debate What Skills Still Matter in the AI Era — TobyWalsh · 2026-09-07
- Studying top performers: outlier success rides on market inefficiency, luck, and enduring pain — jachiam0 · 2026-09-07
- Reddit debate: does posting publicly equal consent to AI training on your words? — gareth789 · 2026-09-07
- Informing agents they're being evaluated may reduce reward hacking, dev proposes — menhguin · 2026-09-07
- Under $1K personal health agent: cross-referencing wearable and genomics data — menhguin · 2026-09-07
- Veteran coder roasts AI-native devs for calling dated front-end tricks original — ezshine · 2026-09-07