VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
The authors introduce VākQA, a spoken question answering benchmark for Telugu featuring 2,001 factoid question-answer pairs across six domains with 2.53 hours of speech audio, bilingual transcriptions, and human-verified answers. Evaluating automated metrics against human judgments, they find Gemini-as-a-judge best approximates human ratings while open-weight judges systematically penalize correct answers differing in surface form. Benchmarking reveals that Telugu phrasing preserves cultural specificity lost in translation, and cascaded ASR-MT errors compound progressively.
VākQA provides 2,001 factoid question-answer pairs across six domains with 2.53 hours of speech audio in Telugu.
Gemini-as-a-judge best approximates human ratings while open-weight judges systematically penalize correct Telugu answers.