Intent-Eval benchmark shows rejected user changes still derail LLM multi-turn task execution

Junle Chen · hf · 2026-10-09

New research uncovers a surprising LLM failure in multi-turn dialogue: when a user proposes a change but rejects it, merely mentioning the rejected change can derail task execution — a "mentioned-as-in-effect" confusion where conversational content is treated as active requirements. Accuracy degradation deepens or persists as interaction continues.

The authors release Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics, to systematically study model behavior under evolving user intent.

They also propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework: a frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training a Student on the full dialogue to follow only requirements that reflect actual user intent.

Original post →

More from Research

Research channel →