PlaylistEval: frontier video-language judges hit only 75.4% accuracy on 100-hour playlists
Shayekh Bin Islam · hf · 2026-09-30
PlaylistEval is an agentic framework building video-language judge benchmarks over 100-hour playlist collections without human annotation, using causally degraded answer pairs that force retrieval. The 630-pair benchmark (7 domains) agrees with humans 93.0% on a 152-pair subset. Evaluating 17 models from 8 families, frontier judges reach only 75.4% pairwise accuracy while open-source judges lag far behind; both retrieval and judgment require multiple modalities, and accuracy degrades as playlists grow.
More from Models
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- DeepSeek is giving users 6 yuan in free API credits via its harness — teortaxesTex · 2026-09-30
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30
- DeepSeek opens community feedback channel: flip a toggle in the harness to help improve its models — teortaxesTex · 2026-09-30
- You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs — heghbalz · 2026-09-30