Microsoft finds LLMs lose track of user intent when conversations evolve
microsoft · hf · 2026-07-24
Static benchmark wins do not carry over when user intent changes
Microsoft Research studies collaborative LLM agents in dynamic conversations, where user intent is revealed, revised, and redirected over multiple turns rather than stated upfront.
The key idea
- They convert static, single-turn tasks into multi-turn evolving-intent conversations.
- The framework preserves the original evaluation protocol, so existing benchmarks can be reused without new annotations.
Main finding
- Across multiple tasks, models that perform strongly in static settings suffer substantial drops when intent evolves.
- The result suggests a fundamental gap: today’s LLMs still do not reliably track and act on changing user intent, even though that skill is critical for future collaborative agents.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11