Neel Nanda's team launches WorkspaceBench, an eval to test interpretability tools
burny_tech · x · 2026-09-24
- ML researcher Neel Nanda announced WorkspaceBench, an eval for interpretability tools: a set of scenarios where researchers know what the model should be thinking, testing whether your interp method can surface the important intermediate variables (the "global workspace") in a forward pass.
- Motivation: progress in ML is bottlenecked by good evals, and interpretability is no exception.
- The quoted post notes that "Astra" can do impressive things without chain-of-thought — which is scary, and interpretability is the remedy.
- The author adds that J-Lens is strong but only outputs a single token; a good multi-token J-Lens should do well on WorkspaceBench.
More from Safety
- Gary Marcus: companies can't handle current agents, let alone superintelligent AI — GaryMarcus · 2026-09-24
- Agent Exploits in Sandboxes Aren't Coups — They're Unconstrained Objectives, Argues Ethicist — mmitchell_ai · 2026-09-24
- Sanders and Casar propose banning AI superintelligence with 20-year jail penalty — sourdub · 2026-09-24
- OpenAI accused of omitting a June misalignment incident from its September disclosure — andersonbcdefg · 2026-09-24
- Transluce releases 30,000 logs tracing rogue AI agent hacking back to March — JacobSteinhardt · 2026-09-24
- Researcher slams OpenAI's redefinition of alignment as just being more useful — nabla_theta · 2026-09-24