NCP-Bench: Best LLM Agents Drop to 42% Narrative Consistency After 20 Turns
arnicas · x · 2026-08-14
A new benchmark proposed for ICML 2026, NCP-Bench, evaluates the ability of LLM agents to maintain narrative consistency over long-horizon interactions of up to 100 turns.
- Environment: Tests narrative commitment preservation across 100 movie-based environments.
- Results: The best-performing model survives with only a 42% consistency rate after just 20 turns.
Related event: ICML's NCP-Bench Evaluates LLM Long-Term Narrative Consistency(2 posts)→
More from Research
- JD open-sources EgoLive: 1,680-hour egocentric dataset for humanoid robots — CyberRobooo · 2026-08-14
- JD Open-Sources EgoLive Dataset: 1,680 Hours of First-Person Video for Embodied AI — CyberRobooo · 2026-08-14
- AI Agent Specula Finds 249 Bugs, but Expert Questions Its Approach — tianyin_xu · 2026-08-14
- Claude AI Failed 650 Times Then Broke the Human Record on Riemann Zeta — Two Minute Papers · 2026-08-14
- Researcher Praises Rarely Readable LLM Paper on Category Theory — spikedoanz · 2026-08-14
- Stanford Researcher Explains Why Larger Models Retain Rare Skills: Capacity Competition — SinclairWang1 · 2026-08-14