african-languages
Filtering by topic african-languages(3)Clear all filters
- PaperarXiv — Language & NLP (cs.CL)23 Jul 2026
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Paul Azunre
DONDO introduces 21 monolingual and 5 multilingual automatic speech recognition base models covering 27 African language varieties, built on the w2v-BERT 2.0 encoder and fine-tuned on read speech from religious texts. The models achieve word error rates of 10–13% across multilingual families, closing the gap to monolingual performance while enabling language-steering via a one-hot prefix. All models are released under Apache-2.0 to support further fine-tuning and commercial use, potentially benefiting over 100 million first-language speakers.
Original abstract
We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily on read speech drawn from religious texts, which offer broad, license-clear and orthographically consistent coverage for languages that otherwise lack transcribed audio. We describe a two-step (and, for one family, three-step) learning-rate-annealed fine-tuning procedure that first adapts a shared multilingual model at a high learning rate and then anneals it to recover, and in several cases surpass, strong monolingual baselines. We further describe a lightweight language-conditioning mechanism that injects a one-hot language identity as a sequence of prefix frames prepended to the acoustic features, allowing a single multilingual checkpoint to be steered to a target language at inference. Across the five multilingual families the annealed models reach average word error rates (WER) of 10-13%, closing most of the gap to monolingual models while covering many languages in a single checkpoint. All models are released on the Hugging Face KhayaAI organisation under the Apache-2.0 license (attribution only) so that others may fine-tune them freely, including for commercial use. We provide a conservative estimate that the languages covered are spoken by on the order of one hundred million first-language speakers, and by substantially more when second-language use is included.
- NewsLanguage Magazine6 Jul 2026
Celebrate World Kiswahili Language Day on July 7
World Kiswahili Language Day will be commemorated on July 7 for the fifth time, highlighting Kiswahili as the first African language to have its own international day. With over 200 million speakers, Kiswahili is one of the most widely spoken languages in Africa and globally.
Original abstract
Tuesday, July 7 will mark the fifth commemoration of World Kiswahili Language Day. Kiswahili is one of the most widely spoken languages in Africa and the world. With over 200 million speakers, Kiswahili is the first and only African language to be granted its own international day. The theme of this year’s celebration is “Kiswahili […] The post Celebrate World Kiswahili Language Day on July 7 appeared first on Language Magazine .
- PaperarXiv — Language & NLP (cs.CL)5 Jul 2026
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
Offiong Bassey Edet, Emmanuel Oyo-Ita, Archibong Okon Archibong, David Effanga Bassey et al.
Presents the first end-to-end text-to-speech study for Efik, a low-resource tonal language in Nigeria, evaluating four neural models and finding that MMS-TTS achieved the highest MOS but tonal errors persisted, highlighting the need for larger corpora and tone-aware modeling.
Original abstract
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.