speech-synthesis
Filtering by topic speech-synthesis(2)Clear all filters
- PaperarXiv — Language & NLP (cs.CL)5 Jul 2026
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
Offiong Bassey Edet, Emmanuel Oyo-Ita, Archibong Okon Archibong, David Effanga Bassey et al.
Presents the first end-to-end text-to-speech study for Efik, a low-resource tonal language, comparing four neural models with a curated corpus. MMS-TTS achieved the highest mean opinion score of 3.80, though tonal errors persisted, highlighting the need for tone-aware modeling.
Original abstract
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.
- PaperarXiv — Language & NLP (cs.CL)18 May 2026
Bridging the Gap: Converting Read Text to Conversational Dialogue
Parshav Singla, Agnik Banerjee, Aaditya Arora, Shruti Aggarwal et al.
A novel approach, Prosodic Adjustment with Conversational Context (PACC), converts read speech into conversational speech using HiFi-GAN, improving naturalness and intelligibility. The method modifies prosodic features like intonation and rhythm, achieving better MOS scores. This technique has applications in virtual assistants and language learning tools.
Original abstract
In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge in this domain is maintaining naturalness and intelligibility while minimizing computational overhead for real-time applications. Traditional read speech often lacks the nuanced prosodic variation essential for natural conversational interactions, posing challenges for applications in virtual assistants, customer service, and language learning tools. This paper introduces a novel approach, Prosodic Adjustment with Conversational Context (PACC), aimed at converting read speech into natural conversational speech used in various modern applications. PACC utilizes advanced deep neural networks to analyze and modify prosodic features such as intonation, stress, and rhythm. Unlike conventional methods, our approach uses High-Fidelity Generative Adversarial Networks (HiFi-GAN) for speech synthesis. Our experimental results demonstrate significant improvements in speech conversion, enhancing naturalness and achieving better model accuracy with additional training on speech datasets. This research establishes new benchmarks in speech conversion tasks and Mean Opinion Score (MOS) evaluation for testing model accuracy, and we show that our approach can be successfully extended to other speech conversion applications.