NeurIPS 2026 to Host Workshop on Rigorous Foundations for LLM Interpretability
hugo_larochelle · x · 2026-08-03
NeurIPS 2026 in Sydney will host the inaugural InterpScience Workshop, aiming to build a more rigorous scientific foundation for large language model (LLM) interpretability.
Key Topics:
- Defining what it means to "understand" an LLM.
- Establishing standards, benchmarks, or evaluation criteria for measurement, causal claims, and falsifiability.
- Drawing lessons for interpretability from neuroscience, statistics, and causal representation learning.
Format:
The workshop will feature an interactive format with multiple breakout sessions led by facilitators from relevant disciplines to encourage cross-disciplinary dialogue.
Important Date:
Papers are due on August 28.
More from Research
- AI Agents Relentlessly Check Math Work, But Fail to Do So in Other Sciences — ChrisGPotts · 2026-08-03
- Adversarial Protocol Review: Using Agents to Catch Flawed Scientific Claims — ChrisGPotts · 2026-08-03
- AI in Theorem Proving: Human Math Abstraction Exceeds Current Tools — prof_g · 2026-08-03
- Unreleased OpenAI model solves 10 major math problems for $2,000 inference cost — Mobile_Distance_9598 · 2026-08-03
- Benchmark Manipulation: You Can Reach Any Conclusion by Controlling Tests — felix_red_panda · 2026-08-03
- Current AI Agents Lack the Human 'Power Law of Practice' — xwang_lk · 2026-08-03