Interpretability researcher: eval and curated-data problems have followed me into interp
benno_krojer · x · 2026-10-02
A researcher who moved from data/eval work to interpretability argues interp now faces the same data-centric evaluation problems: judging a lens of a model's 'inner thoughts' is very hard, with methods like NLA training, LatentLens and steering-vector minimal pairs all depending on carefully curated data. He adds that many strong methods papers carry data contributions too.
Related event: Interpretability Research Hits Data Problems Again(2 posts)→
More from Research
- Anima Anandkumar wants AI to understand physics beyond ChatGPT — nordicinst · 2026-10-02
- Model hit 97% accuracy, ran in production for a year — it was useless — TajyMany · 2026-10-02
- Epoch AI releases ChatGPT usage data sampled from YouGov's US panel — evijit · 2026-10-02
- Workload-aware inference: why batch LLM pipelines should plan queries like databases do — sh_reya · 2026-10-02
- SYNTH paper finds epistemic calibration emerges in models from 300M parameters — cephaloform · 2026-10-02
- Failure Map: 20,168 open Python boundary-case bug repair tasks released — failuremap-f · 2026-10-02