Character.ai says pairwise judges catch video drift better than absolute scores
AI Engineer · youtube · 2026-07-25
- Maor Bril from Character.ai argues that video evaluation cannot rely on single absolute scores: a clip can look polished while still failing on temporal coherence, shot continuity, or story consistency.
- Traditional metrics like CLIP miss these failure modes, and human review does not scale.
- The team found pairwise preference judgments more robust than absolute scoring, and trained a Qwen3-VL judge with Bradley–Terry loss on pairs of real and deliberately broken footage.
- The judge is used as a CI regression gate for every AgentX release at Character.ai, calibrated against human scores, so drift is caught early before users see it.
More from Apps
- Claude Opus 5 prompt tips say old harness habits now waste tokens — tengyanAI · 2026-07-25
- Proof-of-concept wires agents together with Buzz and Lightning bitcoin payments — sull · 2026-07-25
- Agnost AI launches as a PostHog for conversational AI agents — ycombinator · 2026-07-25
- Scape says it raised $3.2M to build an email inbox that works for you — garrytan · 2026-07-25
- Odoo MCP Server lets AI assistants manage ERP workflows in natural language — modelcontextprotocol · 2026-07-25
- Wes Roth spotlights Claude Opus 5, ARC-AGI 3, and Genspark’s $179 SecondBrain Note — Wes Roth · 2026-07-25