Video benchmarks are easy to game: Video-Index keeps only 840 shortcut-proof questions
thoma_gu · x · 2026-10-02
Jiatao Gu's team released Video-Index, a curated meta-benchmark for video understanding, arguing that video benchmarks aren't saturated — they're just easy to game. Using a five-level "attack pyramid" of shortcut attacks, they audited 115 benchmarks and 505K questions:
- On 35 benchmarks, attackers that never see a single frame approach full-video accuracy; on 51 benchmarks with temporal probes, shuffled frames retain a median 96% of full-video accuracy
- Near-duplicate questions make up at least half the items in 63 benchmarks
- The final set: 840 verified hard items (210 per capability group) from 76 sources
On Video-Index, Astra scores 79.3% vs. 55% for humans, while open models score under 20%. Claude Opus 5 outscores the best open-source model by over 37 points, with agent tools adding roughly 20 more. Paper, code, and data are open-sourced.
More from Research
- AMap open-sources ABot-Recon: streaming 3D reconstruction from video with a 12-frame local context — rsasaki0109 · 2026-10-02
- Neuralink pretrains decoders on 50,000+ hours of neural data—dataset may be the real moat — CurieuxExplorer · 2026-10-02
- CyberGym cybersecurity benchmark effectively saturated on verified task subset — aryaman2020 · 2026-10-02
- SemEval-2027 Task 9 calls for teams on 4-language multimodal news framing analysis — preslav_nakov · 2026-10-02
- Alternating prompt and model upgrades lift science agent from 42% to 73% — rohanpaul_ai · 2026-10-02
- ISMIR paper teaches a transformer to play in 12 jazz piano legends' styles — umpedronosapato · 2026-10-02