Why do RL researchers get away with incremental algoslop while interp doesn't?

aryaman2020 · x · 2026-09-07

Pushing back on the claim that interpretability is 'a broadly confused mass of techniques,' aryaman asks whether RL isn't the same — and why its incremental algoslop escapes criticism. In an empirical, actively developed field of science, he argues, this is simply normal.

Related event: Researchers Debate the Value and Boundaries of Mechanistic Interpretability(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →