NARRA-Gym: New Benchmark Finds Frontier LLMs Vary Widely on Long-Horizon Interactive Storytelling
my_cat_can_code · x · 2026-09-26
arXiv paper 2605.08503 introduces NARRA-Gym, an executable environment testing whether LLM agents can sustain a coherent, evolving story across many turns while adapting to a user. It evaluates nine frontier LLMs via controlled LLM-as-judge sweeps over eight personas plus human evaluation, and finds that models producing fluent stories can still fail on robustness, user experience, or resistance-sensitive personalization. The same team also released JobBench, evaluating agents on 130 tasks across 35 occupations, with a NeurIPS appearance planned.
Related event: Bake AI's Two Agent Benchmark Papers Accepted by NeurIPS 2026(2 posts)→
More from Research
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Three papers, one warning: 1% synthetic data can trigger strong model collapse — suchenzang · 2026-09-26
- Simulation beats distillation: the real story of synthetic data is post-training worlds — realsohamparekh · 2026-09-26
- Experts Rise Where LLMs Disagree: rationale labeling cuts codebook revision from months to days — windx0303 · 2026-09-26
- Dev hails continual learning paper: AGI defined in 2000, only now is anyone training for it — willcb · 2026-09-26
- Lawrence Krauss podcast asks whether AI will supercharge scientific paper mills — willcb · 2026-09-26