ICML's NCP-Bench Evaluates LLM Long-Term Narrative Consistency
ICML 2026 introduced NCP-Bench, a new benchmark evaluating LLMs' ability to maintain narrative consistency during long-term interactions of up to 100 turns. Tests reveal that even the strongest models struggle, dropping to a mere 42% consistency after just 20 turns.
2026-08-13 ~ 2026-08-14 · 2 related posts
- Evaluating LLMs: New Benchmark for Long-Horizon Consistency in Interactive Narratives — Yingpeng Ma · 2026-08-13
- NCP-Bench: Best LLM Agents Drop to 42% Narrative Consistency After 20 Turns — arnicas · 2026-08-14