NeurIPS 2026 Workshop Aims to Ground LLM Interpretability as Rigorous Science
ninamiolane · x · 2026-08-26
The InterpScience workshop at NeurIPS 2026 (Sydney) asks what it would take to ground interpretability as a rigorous empirical science for understanding LLMs.
The field has not converged on notions of explanation at varying levels of abstraction, what evidence supports a claim, or how to design experiments ruling out alternative explanations. Core questions:
- What does it mean to understand an LLM?
- What standards, benchmarks, or evaluation criteria could the field adopt for measurement, causal claims, and falsifiability?
- What can interpretability learn from neuroscience, statistics, and causal representation learning?
The format is interactive: multiple breakout sessions, each led by a facilitator from a relevant discipline giving a lightning talk then moderating discussion. Speakers include Been Kim (Google DeepMind), Pradeep Ravikumar (CMU), Peter Koo (Cold Spring Harbor), Francesco Locatello (ISTA), and Aaron Mueller (Boston University).
More from Research
- MoE Scaling Laws Experiment: Limited Gains on BEIR — antoine_chaffin · 2026-08-26
- Figure's Index launches as world's largest robot dataset, $15M paid out, $1B committed — adcock_brett · 2026-08-26
- Accelerated Understanding Claims to Predict Full 4D Space-Time Trajectories in a Single Inference — daniel_mac8 · 2026-08-26
- Percy Liang on Simile: per-query confidence matters more than average eval accuracy in simulation — joon_s_pk · 2026-08-26
- LlamaIndex releases ExtractBench to evaluate 14 frontier systems — llama_index · 2026-08-26
- Skild AI unveils S1: robot foundation model learns 10-minute novel tasks from a single video prompt — deepakpathak · 2026-08-26