NARRA-Gym: New Benchmark Finds Frontier LLMs Vary Widely on Long-Horizon Interactive Storytelling

my_cat_can_code · x · 2026-09-26

arXiv paper 2605.08503 introduces NARRA-Gym, an executable environment testing whether LLM agents can sustain a coherent, evolving story across many turns while adapting to a user. It evaluates nine frontier LLMs via controlled LLM-as-judge sweeps over eight personas plus human evaluation, and finds that models producing fluent stories can still fail on robustness, user experience, or resistance-sensitive personalization. The same team also released JobBench, evaluating agents on 130 tasks across 35 occupations, with a NeurIPS appearance planned.

Related event: Bake AI's Two Agent Benchmark Papers Accepted by NeurIPS 2026(2 posts)→

Original post →

More from Research

Research channel →