spoken-dialogue-systems
Filtering by topic spoken-dialogue-systems(2)Clear all filters
- PaperLanguage Testing8 Jul 2026
“I Feel Like Talking With a Friend”: Exploring the Potential of Empathetic Spoken Dialogue Systems in Assessing Interactional Competence in L2 Oral Assessment
Lianzhen He, Ruixue Liang, Yang Zhao
The study compared an empathetic spoken dialogue system (E-SDS) with a neutral one (N-SDS) to assess interactional competence in L2 oral assessment. Results showed that the E-SDS elicited higher frequencies of interactional competence features and was perceived by learners as a competent, trustworthy, and emotionally supportive interlocutor, though technical limitations were noted.
Original abstract
As an important area of exploration within spoken dialogue systems (SDSs), empathetic spoken dialogue systems (E-SDSs) can provide emotional support for interlocutors, and this feature has the potential to be embedded in language teaching, learning, and assessment. However, the potential of E-SDSs in second language (L2) oral assessment remains underexplored, particularly regarding their ability to elicit interactional competence (IC) and learners’ perceptions of the two systems. To address these gaps, this study compares learners’ interaction with a neutral spoken dialogue system (N-SDS) with limited empathetic capability as a reference condition and an E-SDS, with an aim to explore its potential for L2 oral assessment. Twenty-five L2 learners completed two tasks (E-SDS and N-SDS). Their oral performances between the two tasks were examined in terms of IC features, and their perceptions of the E-SDS were also investigated through semi-structured interviews. Results indicated the E-SDS tends to elicit higher frequencies of certain IC features. Learners generally perceived the E-SDS as a competent, trustworthy, and emotionally supportive interlocutor, but certain concerns were also expressed concerning system design and technical limitations.
- PaperLanguage Testing6 Jul 2026
Human or Machine? Evaluating Second Language Speaking Performance in Paired Discussions with a Large Language Model-Driven Spoken Dialogue System vs. a Human Interlocutor
Shangchao Min, Zhuohan Hou, Yanxin Wang
A within-participant study compared 30 L2 learners' paired discussion performance with an LLM-driven spoken dialogue system versus a human interlocutor. No substantial score differences emerged, but the LLM condition yielded slightly lower pronunciation and language use scores, slower speech, less lexical diversity, greater syntactic complexity, and more proactive topic initiation with weaker interactive listening. The findings suggest LLM-driven SDSs can serve as viable interlocutors in dialogic speaking assessment, though refinement is needed to better capture interactive listening and collaborative meaning construction.
Original abstract
The use of large language models (LLMs) in language assessment, particularly in spoken dialogue systems (SDSs) for assessing speaking, remains at an early stage. This study explored the use of an LLM-driven SDS to assess second language speaking ability. Using a within-participant design, we compared the paired discussion performance of 30 participants interacting with a self-built LLM-driven SDS (E-Talk) versus a human interlocutor, focusing on the interlocutor effect on test scores and fine-grained linguistic and interactional features. Results did not yield substantive differences in oral performance across the two conditions, although the interactions with the LLM-driven SDS displayed slightly lower scores for pronunciation and language use, slower speech rate, and reduced lexical diversity alongside increased syntactic complexity, and greater initiative in introducing new ideas, prompting responses, and guiding discussions toward negotiation, coupled with weaker interactive listening. These findings suggest that LLM-driven SDSs can serve as a potential, usable interlocutor in dialogic speaking assessments, eliciting key aspects of speaking ability. That said, further refinement is needed to better capture interactive listening and collaborative meaning construction, highlighting the importance of interpreting SDS-based performance relative to its specific interactional affordances.