IntentFlux Benchmarks 'Intent Drift' in LLM Agents: Scores Fall from 0.476 to 0.384 as Users Change Their Minds

Yanjie Zhang · hf · 2026-10-02

Researchers introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes to study 'intent drift'—when superseded parts of a user's intent keep influencing an LLM agent's final answer or tool action. Across 627 calibration cases, mean task score falls from 0.476 to 0.384 as superseded and withdrawn information accumulates, and eight models all score significantly lower recovering the final task from an evolving dialogue than receiving it in one turn. The accompanying StateForge explicitly maintains active requirements before generation, lifting General-Test mean score from 0.367 to 0.467; even ground-truth final states don't recover single-turn performance, showing state-estimation error explains only part of the gap.

Original post →

More from coding & agent

coding & agent channel →