ai
HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models
arXiv: Computers and SocietyInternationalHigh confidence1 min
What changed
A new benchmark, HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), has been introduced to assess bias in audio language models (Audio-LLMs). Comprising 87,000 real human audio samples from 843 demographically diverse participants, HEAR facilitates comprehensive evaluation through both Multiple Choice Question Answering (MCQA) and open-ended tasks. Initial evaluations reveal that voice-conditioned bias is model-specific and that personalization instructions consistently worsen demographic disparities within these models.
Why it matters
The introduction of a demographically diverse benchmark for Audio-LLMs is critical for understanding and mitigating algorithmic bias in voice-enabled technologies. Identifying that bias is model-specific and worsened by personalization highlights significant risks to equitable service delivery and user trust across various applications.
What to watch
HEAR is a new, large-scale, ecologically valid benchmark with 87,000 real human audio samples from 843 diverse participants.
Forward consideration, not a verified fact.
Reported by arXiv: Computers and Society, International. The document itself is not reproduced here.
Read the original publication