AfriBED: a Speech Recognition Bias Evaluation Dataset for Low Resource African Languages

Sikasote, Claytone and Suleman, Hussein and Buys, Jan (2026) AfriBED: a Speech Recognition Bias Evaluation Dataset for Low Resource African Languages.

Full text not available from this repository. (Use alternate locations listed below)

Abstract

Despite advancements in automatic speech recognition systems (ASR), current systems exhibit speaker-attribute bias, i.e., their speech recognition performance is biased against certain speakers based on their demographic attributes such as gender, accents, and speaker type. However, accurately assessing speaker attribute bias in a given ASR system requires the availability of a bias evaluation dataset with comparable speech across subgroups of a speaker attribute. In this paper, we introduce African speech Bias Evaluation Dataset (AfriBED), a novel speech recognition bias evaluation dataset for two low resource African languages, Bemba and Nyanja, to detect and estimate gender and speaker-type bias for ASR systems. The dataset comprises 9 hours of speech with 3928 utterances recorded by 80 speakers for Bemba, while the Nyanja set includes 6260 utterances with approximately 18 hours of speech recorded by 127 speakers. The results of our baseline experiments demonstrate that the datasets are useful for examining gender and speaker-type attribute bias in the target languages.

Item Type: Preprint
Uncontrolled Keywords: ASR, Fairness, Bemba, Nyanja, Africa language
Subjects: Computing methodologies > Artificial intelligence > Natural language processing
Computing methodologies > Artificial intelligence > Natural language processing > Speech recognition
Date Deposited: 05 Oct 2026 10:19
Last Modified: 05 Oct 2026 10:19
URI: https://pubs.cs.uct.ac.za/id/eprint/1803

Actions (login required)

View Item View Item