CineSubBench: 1,012 films and 8.13M subtitle entries test LLMs on long-form narrative and culture
Mir Tafseer Nayeem · hf · 2026-09-30
A new benchmark, CineSubBench, evaluates LLM long-context film understanding from multilingual movie subtitles—filling a gap in domains beyond law, medicine, and code.
- Scale: 1,012 films with complete subtitles in 6 languages—6,072 tracks and 8.13M timestamped entries; models must reconstruct characters, relationships, events, causality, and themes from thousands of short utterances.
- Tasks: 7 tasks spanning narrative reconstruction/abstraction, genre prediction, age suitability, country-specific ratings across 10 national systems, and subtitle-grounded language safety—a matched multilingual, multicultural (MultiX) setting.
- Findings across 9 LLMs: plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies widely by model and language; national rating systems expose distinct calibration patterns; strong profanity is far easier to ground than mild obscenity.
CineSubBench establishes film as a long-context LLM evaluation domain.
More from Research
- Researcher predicts AI labs will soon pivot from math conjectures to materials and drug discovery — tak3sh8 · 2026-09-30
- Video lecture series by Stephen Wright, Yousef Saad and Peter Bartlett on ML optimization now available — caglar_ee · 2026-09-30
- Arbor: open-source framework for AI agents doing autonomous long-horizon research — burkov · 2026-09-30
- Physics-aware losses keep grain boundaries real when AI generates alloy microstructures — bravo_abad · 2026-09-30
- SOSP26 Paper YoloFS Targets Agent Filesystem Misuse, Built From 290 Real Incident Reports — tianyin_xu · 2026-09-30
- Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds — maksym_andr · 2026-09-30