CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 31 upvotes

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

QUESTION — How do proprietary and open-weight models perform on spoken factoid question answering in Telugu, and how reliable are automatic evaluation methods for this low-resource setting?

The authors introduce VākQA, a spoken question answering benchmark for Telugu featuring 2,001 factoid question-answer pairs across six domains with 2.53 hours of speech audio, bilingual transcriptions, and human-verified answers. Evaluating automated metrics against human judgments, they find Gemini-as-a-judge best approximates human ratings while open-weight judges systematically penalize correct answers differing in surface form. Benchmarking reveals that Telugu phrasing preserves cultural specificity lost in translation, and cascaded ASR-MT errors compound progressively.

VākQA provides 2,001 factoid question-answer pairs across six domains with 2.53 hours of speech audio in Telugu.

Gemini-as-a-judge best approximates human ratings while open-weight judges systematically penalize correct Telugu answers.

Bhavanaakkiraju · 17 Sept 2026 read the original ↗
↑