EMNLP '26 Paper: KnowSim Evaluates LLM Information Calibration
QVeraLiao · x · 2026-08-26
Paper accepted to EMNLP 2026 introduces KnowSim, a framework with a user simulator that tracks knowledge levels to evaluate LLMs.
- Problem: Current simulators ignore user knowledge states, failing to assess if AI tailors info to novices vs. experts.
- Solution: KnowSim maintains a knowledge state graph grounded in cognitive science, computing metrics for Knowledge Gain, Delivery Calibration, and Cognitive Overload.
- Results: Rankings align with human judgments (73-74%) across 705 sessions. Testing on 9 LLMs reveals the 'best' model shifts by user knowledge level, uncovering interactions invisible to standard evals.
More from Research
- NeurIPS mechanistic interpretability workshop extends deadline — ninamiolane · 2026-08-26
- NeurIPS Findings track deadline extended to Sept 7th — ninamiolane · 2026-08-26
- Nonprofit Sophron Research Launches to Develop AI Model Evaluations — ryan_t_lowe · 2026-08-26
- NeurIPS workshop CFP: Interpreting Agent Behavior — mdredze · 2026-08-26
- Reasoning Models Outperform via Higher Recovery Rates, Not Just "More Thinking" — Jeande_d · 2026-08-26
- Study Finds Reasoning Models' Amplified Behaviors Weakly Linked to Correctness — Jeande_d · 2026-08-26