text-to-speech
Filtering by topic text-to-speech(3)Clear all filters
- PaperarXiv — AI in Education (cs.CY)14 Jul 2026
A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential
Gendo Kumoi, Fumie Watanabe, Tota Suko, Takashi Ishida et al.
Proposes a semi-automated system combining LLMs and TTS to generate dialogue-based lessons for high school students. A quasi-experiment with 245 students found that dialogue TTS improved comprehension and cognitive engagement over single-speaker TTS, though single TTS had better audio naturalness. The results support the educational potential of TTS audio and dialogue formats in language learning.
Original abstract
This study proposes a semi-automated system for generating dialogue-based lessons using Large Language Models (LLMs) and Text-to-Speech (TTS) technology, and exploratorily examines its educational potential via a practical quasi-experiment. The system augments rather than replaces educators through a three-stage human-in-the-loop workflow (LLM-based slide/narration generation, educator review, automated audiovisual integration), and introduces a novel method for generating Expert-Novice dialogue narration based on cognitive apprenticeship theory. In a study of 245 first-year high school students who sequentially experienced three lesson formats (instructor voice, single-speaker TTS, dialogue TTS; content differed across sessions, limiting format/content separation), we conducted within-subject (Friedman test, N<=183) and repeated cross-sectional (Mann-Whitney U, N=229/206) analyses. TTS audio did not substantially degrade the learning experience versus instructor voice, supported by TOST equivalence testing. Dialogue TTS was significantly superior to single TTS in comprehension (p=.006, q=.025) and cognitive engagement (p=.019, q=.048); enjoyment was non-significant after FDR correction (q=.081) but reached significance after controlling for prior knowledge (proportional-odds model, OR=1.65, q=.025), and these advantages were not attributable to prior-knowledge imbalance. Conversely, single TTS was superior in audio naturalness (p<.001, q<.001, r=-.238), revealing a trade-off between dialogue's benefits and higher extraneous cognitive load. Dialogue format was preferred by 66.9% of learners as most enjoyable (p<.001). These results reflect a fixed-order design; replication is needed before generalizing them as effects of lesson format. This study provides a theoretical and empirical basis for the educational acceptability of TTS audio and for TTS lesson-format design.
- PaperarXiv — Language & NLP (cs.CL)5 Jul 2026
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
Offiong Bassey Edet, Emmanuel Oyo-Ita, Archibong Okon Archibong, David Effanga Bassey et al.
Presents the first end-to-end text-to-speech study for Efik, a low-resource tonal language in Nigeria, evaluating four neural models and finding that MMS-TTS achieved the highest MOS but tonal errors persisted, highlighting the need for larger corpora and tone-aware modeling.
Original abstract
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.
- PaperReCALL1 Dec 2025
Examining the effectiveness of integrating corpus-based and AI approaches for English speaking practice
Hsueh Chu Chen, Xiaona Zhou, Jing Xuan Tian
An online English speaking training approach integrating a self-developed spoken corpus, generative AI, and text-to-speech tools was developed and evaluated. Pre- and post-test results showed improvements in participants' speaking performances, including increased use of complex sentences and fewer vowel errors. Participants reported positive attitudes and highlighted the benefits of combining corpora and AI for accurate feedback and interactive learning.
Original abstract
This study developed and evaluated an online English speaking training approach that integrates corpora and artificial intelligence (AI) tools. The training integrated a self-developed spoken corpus, generative AI tools, and text-to-speech AI tools. Pre- and post-test results identified improvements in participants’ speaking performances. Participants attempted to use more positive linguistic features (e.g. producing complex sentences more frequently) and avoid using negative linguistic features (e.g. reducing the number of vowel errors) after receiving the training. Participants showed positive attitudes towards this corpus-based and AI-integrated English oral ability learning approach and affirmed the importance of integrating both tools. The corpus helped raise participants’ awareness of features that influence speaking performance and offered prompt engineering and feedback-checking functions, while the generative AI tools provided useful feedback and tailor-made sample responses. Additionally, text-to-speech AI tools offered learners with tailor-made native speaker samples for imitation and helped learners learn pausing. Results also revealed that this approach helped create an interactive oral ability learning environment, and the combination of corpora and AI tools provided more accurate feedback for each subskill of speaking.