VākQA: A Telugu spoken QA benchmark exposing flaws in LLM-as-judge evaluation

SPL-IIITH · hf · 2026-09-18

Researchers released VākQA, the first spoken factoid QA benchmark for Telugu: 2,001 QA pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human-verified references.

Judge validation found Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers differing in surface form from references.

Benchmarking proprietary and open-weight models revealed that Telugu phrasing carries cultural specificity lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. The benchmark is publicly released.

Original post →

More from Research

Research channel →