Evaluating LLMs: New Benchmark for Long-Horizon Consistency in Interactive Narratives

Yingpeng Ma · hf · 2026-08-13

This study introduces a new benchmark to evaluate the performance of large language models in interactive narratives. It formalizes the concept of 'narrative commitment preservation' to specifically assess long-horizon logical consistency in interactive storytelling with LLMs.

Related event: ICML's NCP-Bench Evaluates LLM Long-Term Narrative Consistency(2 posts)→

Original post →

More from Research

Research channel →