Intent-Eval benchmark shows rejected user changes still derail LLM multi-turn task execution
Junle Chen · hf · 2026-10-09
New research uncovers a surprising LLM failure in multi-turn dialogue: when a user proposes a change but rejects it, merely mentioning the rejected change can derail task execution — a "mentioned-as-in-effect" confusion where conversational content is treated as active requirements. Accuracy degradation deepens or persists as interaction continues.
The authors release Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics, to systematically study model behavior under evolving user intent.
They also propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework: a frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training a Student on the full dialogue to follow only requirements that reflect actual user intent.
More from Research
- Physics Benchmark for Video World Models: Best of 8 SOTA Scores Just 57.76/100 — nikola_mr64990 · 2026-10-09
- AI scales worse with more tokens than humans do with more time — but is catching up — tobyordoxford · 2026-10-09
- IIT Madras professor cuts privacy-preserving LLM calls from 100+ to just 4 — ravi_iitm · 2026-10-09
- DeskForge-1M: a 1M+ sample GUI dataset with grounding labels trends on Hugging Face — docling-project · 2026-10-09
- Samsung's reViT: one recurrent Transformer block matches full-depth encoders with ~70% fewer parameters — SamsungResearch · 2026-10-09
- AgenticBBO-Bench benchmarks LLM agents for black-box optimization; GPT-6 Astra and DeepSeek-V4.1-Flash on Pareto frontier — Ming Chen · 2026-10-09