Microsoft finds LLMs lose track of user intent when conversations evolve
microsoft · hf · 2026-07-24
Static benchmark wins do not carry over when user intent changes
Microsoft Research studies collaborative LLM agents in dynamic conversations, where user intent is revealed, revised, and redirected over multiple turns rather than stated upfront.
The key idea
- They convert static, single-turn tasks into multi-turn evolving-intent conversations.
- The framework preserves the original evaluation protocol, so existing benchmarks can be reused without new annotations.
Main finding
- Across multiple tasks, models that perform strongly in static settings suffer substantial drops when intent evolves.
- The result suggests a fundamental gap: today’s LLMs still do not reliably track and act on changing user intent, even though that skill is critical for future collaborative agents.
More from Research
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27
- TechCrunch says brain-wave signals could be the next unlock for physical AI training — TechCrunch AI · 2026-07-27