MINTEval, accepted at NeurIPS, shows LLM agents fail at tracking evolving contexts
EliasEskin · x · 2026-10-02
MINTEval, a benchmark accepted to NeurIPS, tests LLM agents and memory systems in long-horizon, continually changing environments (Wikipedia pages, Git repos). Key findings:
- Avg. 86 context updates with interference; 5 question types incl. long-range lookback and multi-target reasoning
- 4 realistic domains: state tracking, multi-turn dialogue, Wikipedia revisions, GitHub commits
- Avg. 138.8k tokens per instance (up to 1.8M); 95.6% human-verified QAs
Across 7 models, agents struggle to track changes, recall precise details, and combine conflicting updates: models confuse old vs. new information, memory systems insert redundant entries, and retrieval often fails—rooted in limitations of retrieval and memory construction.
More from coding & agent
- Workday opens its developer platform to indie devs for building AI agents on enterprise workflows — HeyNayeem · 2026-10-02
- One prompt, 8h 21min: Claude produces an 8-minute documentary, 96 takes, fully autonomous — GCWebDesigner · 2026-10-02
- After going all-in on agents, 9 of 11 sprint tickets are specced by PMs and built by AI — trvklhn666 · 2026-10-02
- ComfyUI Ships Comfy Agent to All Comfy Cloud Users: AI Plans, Builds and Fixes Workflows — cpaik · 2026-10-02
- Building a personal agent harness: the workflow matters more than the model — ghumare64 · 2026-10-02
- Surrendering to Claude Code's permission prompts: granting Bash(*) to all — Daniel_Farinax · 2026-10-02