VākQA: A Telugu spoken QA benchmark exposing flaws in LLM-as-judge evaluation
SPL-IIITH · hf · 2026-09-18
Researchers released VākQA, the first spoken factoid QA benchmark for Telugu: 2,001 QA pairs across six domains, 2.53 hours of speech audio, bilingual transcriptions, and human-verified references.
Judge validation found Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers differing in surface form from references.
Benchmarking proprietary and open-weight models revealed that Telugu phrasing carries cultural specificity lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. The benchmark is publicly released.
More from Research
- Anthropic: Claude-optimized bio models run 4x faster with 100x less GPU time — BenBlaiszik · 2026-09-18
- Empirical study dissects which coding-agent harness components actually help — HamedZamani · 2026-09-18
- Q-Planning adds a Q-function to policies like pi-0.5 so robots improve without interventions — micoolcho · 2026-09-18
- NVIDIA-backed open source AI models children's hearts in seconds, replacing 6-hour manual workflow — nvidia · 2026-09-18
- Michigan Aero video showcases generative AI for aerodynamic design and turbulence control — ricardovinuesa · 2026-09-18
- He Kaiming's team proposes ELF: continuous embedding-space diffusion LMs beat discrete DLMs with fewer steps — alec_helbling · 2026-09-18