TimelineBench: Best of 16 AI Agents Passes Just 26.8% of 56 Real Video-Editing Tasks
ycombinator · x · 2026-10-02
A new arXiv paper introduces Timeline-Bench, a benchmark of 56 real video-editing tasks where agents must turn raw production footage into a finished cut — selecting dialogue takes, shaping interviews into stories, or cutting commercials from product shots, voiceovers and graphics. Each task ships with a brief, source assets, a container, and tests; quality checks are calibrated on 2,582 blind judgments from 43 video editors.
Key findings:
- 16 agents pairing frontier models with harnesses like Codex, Claude Code and OpenCode were evaluated
- The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of 56 tasks (26.8%); the average agent resolves 14.0%
- Human editors prefer the reference edit in 83.5% of judgments
- Of 771 failed runs, 562 fail only the quality test: agents perceive footage via stills and transcripts and check renders for defects, not craft
Tasks, verifier and per-run results are released openly.
More from coding & agent
- We shipped 2,000 MCP integrations. Enterprise customers wanted only their own 20 — alexcovo_eth · 2026-10-02
- "The repos are now the prompts, 100%": Ofir Press on AI coding's shifting context — OfirPress · 2026-10-02
- OpenHands announces SOC 2 Type II compliance for enterprise agent security — rajistics · 2026-10-02
- Grady Booch: Coding agents write code well, but architecture still needs human judgment — Grady_Booch · 2026-10-02
- Jev pitches 'decision primitives': models plugging into logic without text — hardimanjames · 2026-10-02
- LangChain's Sproul: agent core pattern unchanged for a year, "we've been at AGI for four months" — BraceSproul · 2026-10-02