← Back to brief
ResearchOfficialPreprintarXiv Audio and Speech Processing

DCASE 2026 Task 5 Launches Audio-Dependent Question Answering Benchmark

DCASE 2026 Task 5 introduces Audio-Dependent Question Answering (ADQA), a new benchmark designed to evaluate whether large audio-language models answer questions based on audio content rather than relying on textual priors. The ADQA-Bench evaluation set contains 3,000 items across music, speech, and environmental audio, filtered to remove questions solvable from text alone. In its inaugural run, the top system achieved 58.33% accuracy, with evaluation accuracy dropping by an average of 11.91 percentage points compared to development scores.

Why it matters: This benchmark provides a more rigorous and targeted evaluation of audio-language models by ensuring that performance reflects genuine audio understanding rather than exploitation of textual shortcuts.

Full story at: arXiv Audio and Speech Processing