Study: question order affects AI diagnostic accuracy; final accuracy doesn't measure info seeking
marinkazitnik · x · 2026-08-19
A study by Harvard Medical School, Broad Institute, Kempner Institute, and Google DeepMind formalizes multi-turn information seeking as a k-underspecified constraint satisfaction problem, introducing MT-INFOSEEK benchmark with 5,251 problems and 9,006 tasks across math, logic, gene regulation, clinical decision pathways, and 20 Questions. Key findings: models detect missing info; question order matters—wrong first question reduces final accuracy even if all variables are eventually obtained; final accuracy doesn't measure information seeking—models may produce plausible answers without sufficient info.
More from Research
- Claude autonomously designs protein binders with 93% success rate — ycombinator · 2026-08-19
- Classic Paper: The Surprising Creativity of Digital Evolution — Ghost_Pilot_MD · 2026-08-19
- HarmProfile Benchmark: Harmfulness and Diversity Rise with Model Capability — Zhouyuan Ma · 2026-08-19
- Why Removing the Vision Encoder Can Be Better — From an Infra Perspective — liuziwei7 · 2026-08-19
- Matmul Optimization Bottleneck: Data Movement, Not Multiplications — yaroslavvb · 2026-08-19
- Hot take: "Skill optimization" is just repackaged prompt optimization — IanArawjo · 2026-08-19