PlaylistEval: frontier video-language judges hit only 75.4% accuracy on 100-hour playlists

Shayekh Bin Islam · hf · 2026-09-30

PlaylistEval is an agentic framework building video-language judge benchmarks over 100-hour playlist collections without human annotation, using causally degraded answer pairs that force retrieval. The 630-pair benchmark (7 domains) agrees with humans 93.0% on a 152-pair subset. Evaluating 17 models from 8 families, frontier judges reach only 75.4% pairwise accuracy while open-source judges lag far behind; both retrieval and judgment require multiple modalities, and accuracy degrades as playlists grow.

Original post →

More from Models

Models channel →