interlocutor-effect
Filtering by topic interlocutor-effect(1)Clear all filters
- PaperLanguage Testing6 Jul 2026
Human or Machine? Evaluating Second Language Speaking Performance in Paired Discussions with a Large Language Model-Driven Spoken Dialogue System vs. a Human Interlocutor
Shangchao Min, Zhuohan Hou, Yanxin Wang
A study compared second language speaking performance in paired discussions with an LLM-driven spoken dialogue system versus a human interlocutor. No significant overall differences were found, but the LLM condition showed slightly lower pronunciation and language use scores, slower speech rate, reduced lexical diversity yet higher syntactic complexity, and more initiative in discussion management. The findings suggest LLM-driven SDSs can serve as usable interlocutors but require refinement to better capture interactive listening.
Original abstract
The use of large language models (LLMs) in language assessment, particularly in spoken dialogue systems (SDSs) for assessing speaking, remains at an early stage. This study explored the use of an LLM-driven SDS to assess second language speaking ability. Using a within-participant design, we compared the paired discussion performance of 30 participants interacting with a self-built LLM-driven SDS (E-Talk) versus a human interlocutor, focusing on the interlocutor effect on test scores and fine-grained linguistic and interactional features. Results did not yield substantive differences in oral performance across the two conditions, although the interactions with the LLM-driven SDS displayed slightly lower scores for pronunciation and language use, slower speech rate, and reduced lexical diversity alongside increased syntactic complexity, and greater initiative in introducing new ideas, prompting responses, and guiding discussions toward negotiation, coupled with weaker interactive listening. These findings suggest that LLM-driven SDSs can serve as a potential, usable interlocutor in dialogic speaking assessments, eliciting key aspects of speaking ability. That said, further refinement is needed to better capture interactive listening and collaborative meaning construction, highlighting the importance of interpreting SDS-based performance relative to its specific interactional affordances.