IntentFlux Benchmarks 'Intent Drift' in LLM Agents: Scores Fall from 0.476 to 0.384 as Users Change Their Minds
Yanjie Zhang · hf · 2026-10-02
Researchers introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes to study 'intent drift'—when superseded parts of a user's intent keep influencing an LLM agent's final answer or tool action. Across 627 calibration cases, mean task score falls from 0.476 to 0.384 as superseded and withdrawn information accumulates, and eight models all score significantly lower recovering the final task from an evolving dialogue than receiving it in one turn. The accompanying StateForge explicitly maintains active requirements before generation, lifting General-Test mean score from 0.367 to 0.467; even ground-truth final states don't recover single-turn performance, showing state-estimation error explains only part of the gap.
More from coding & agent
- Open-source Harness fork moves coding agents out of the app into orca — dee_hw · 2026-10-02
- LAHacks build Residue uses acoustic analysis and AI agents to personalize your study environment — jonmarkgo · 2026-10-02
- awesome-jev indexes 700 production tools around TypeSafe AI's decision model Jev — Remarkable-Gur719 · 2026-10-02
- $10k of AI inference ports TS to C++ in days, a job for expert teams over years — kristoph · 2026-10-02
- Basis Theory launches revocable credentials letting AI agents act without touching your secrets — km · 2026-10-02
- smolvm v1.22 ships near-instant VM resume for undoing agent actions, 6.5k stars — LoganGrasby · 2026-10-02