speaking-assessment
Filtering by topic speaking-assessment(5)Clear all filters
- PaperLanguage Testing24 Jul 2026
Investigating the Real-World Relevance of an Academic English Speaking Test: Extrapolating Subjective Evaluations and Linguistic Performance Characteristics
Daniel R. Isbell, Dustin Crowther, Jieun Kim, Yoonseo Kim
The study examined correlations between TOEFL Essentials speaking scores and linguistic characteristics with academic speaking tasks in lab and course settings. Strong correlations were found, particularly for fluency and accuracy, supporting the extrapolation of test performances to academic contexts.
Original abstract
To support use of tests in academic contexts, it is critical to demonstrate that test scores and test performances are associated with performance in academic settings—an inferential link referred to as extrapolation in argument-based validation frameworks. TOEFL Essentials is a newer test designed to measure both general and academic English and is intended for use in higher education. The TOEFL Essentials speaking section consists of Virtual Interview, Read Aloud, and Listen & Repeat tasks, the latter two of which elicit highly constrained responses that may be less reflective of academic speaking tasks. In this study, we examined correlations of scores and linguistic characteristics across TOEFL Essentials speaking performances and (a) lab-based academic tasks (graph description, lecture response) for 149 students and (b) an authentic course-based speaking task for 65 students. Strong correlations (.65 < r < .80) were found between TOEFL Essentials speaking scores and evaluations of academic speaking. Among linguistic characteristics, fluency and accuracy variables demonstrated the largest and most consistent correlations across test and non-test tasks. Findings provide evidence relevant to the extrapolation of TOEFL Essentials speaking performances, which are based in part on highly constrained tasks, to academic settings and help inform decisions about test use.
- PaperOpenAlex — TESOL research21 Jul 2026
The Impact of AI Empowerment on the Development of College Students' English Speaking Skills: A Mixed-Methods Analysis Based on Speech Competence Assessment and Self-Directed Learning Efficiency
Jian Li, Luo ZhiYao
This study found that AI empowerment significantly improved Chinese college students' oral English fluency, pronunciation, and lexical resources through a six-week quasi-experiment using the Doubao AI platform. It also revealed a synergistic effect between self-determination theory and self-regulated learning that enhanced self-directed learning efficiency, though a 'situational gap' between virtual practice and real interaction was identified.
Original abstract
In the context of the digital transformation of global education, artificial intelligence (AI) has become a transformative force in the field of second language acquisition. This study explores the impact of AI empowerment on the development of Chinese college students’ oral English competence and the efficiency of self-directed learning (SDL). The core of the research is to explore the synergy between the self-determination theory (SDT) and the self-regulated learning (SRL) model to clarify the interaction between technology empowerment and psychological mechanisms. Under this framework, SDL is the main learning mode, while SDT provides motivational basis (answering students “why” initiates SDL), and SRL provides strategic guarantee (answering students “how” manages learning process). This study adopts a mixed research method and conducts a six-week quasi-experiment with 20 International Business English majors students from Guangdong University of Foreign Studies. During the winter vacation, participants used the Doubao AI platform for autonomous speaking practice and received instant, data-driven feedback. The data were comprehensively analyzed through pre-test and post-test (measuring fluency, lexical resources, accuracy and pronunciation), validated questionnaires and qualitative interviews. The results show that AI empowerment significantly improves objective speaking ability, especially in terms of fluency, pronunciation and lexical resource. More importantly, the study found that there is a deep synergistic effect between the learner’s motivation process and the strategy process: the satisfaction of the two basic psychological needs of autonomy and ability (SDT) provides the internal motivation for students to enter the three cycle stages of SRL; the effective strategy execution guided by the role of AI “virtual supervisor” further strengthens the learners’ sense of achievement, thus jointly improving the overall SDL efficiency. Despite these advances, the study also found a “situational gap” between virtual practice and real social interaction. This study highlights the importance of an AI-supported teaching model that combines technical efficiency with human-led strategic guidance to optimize the self-directed development of oral English in the digital age.
- PaperLanguage Testing8 Jul 2026
Assessing Interactional Competence Through Generative AI: Comparing Large Language Models as AI Interlocutors in the Paired Oral Discussion Test
Inyoung Na
This study compared GPT-4o and Claude 3.5 Sonnet as AI interlocutors in paired oral discussion tests for assessing interactional competence. The results showed that Claude outperformed GPT-4o in eliciting interactional competence features and was perceived as more authentic by test takers, highlighting the need for construct-driven evaluation criteria when selecting LLMs for language assessment.
Original abstract
Interactional competence (IC) is essential for oral communication assessment, yet human partner variability can introduce construct-irrelevant variance in paired speaking tests. As an alternative to a test with a human interlocutor, this study describes the development of a large language model (LLM)-driven Spoken Dialogue System and compares GPT-4o to Claude 3.5 Sonnet to inform model selection for IC assessment. Twelve international students completed paired discussion tasks with both LLMs in counterbalanced order. System performance was evaluated through breakdown activation consistency, stance maintenance, and persona adherence. Test-taker performances were analyzed using interactional discourse analysis to identify IC features across three dimensions: topic management, interactional management, and interactive listening. Semi-structured interviews explored test takers’ perceptions of the AI partners. Results showed Claude outperformed GPT-4o in eliciting IC features, successfully activating communication breakdown strategies and maintaining oppositional stance, thereby creating more opportunities for test takers to demonstrate key IC abilities. Test takers perceived Claude as more authentic and natural, while GPT was perceived as more artificial. These findings demonstrate that different LLMs create distinct interactional conditions affecting both IC elicitation and test-taker perceptions. The findings highlight the need for construct-driven evaluation criteria when selecting LLMs for language-assessment contexts.
- PaperLanguage Testing6 Jul 2026
Human or Machine? Evaluating Second Language Speaking Performance in Paired Discussions with a Large Language Model-Driven Spoken Dialogue System vs. a Human Interlocutor
Shangchao Min, Zhuohan Hou, Yanxin Wang
A study compared second language speaking performance in paired discussions with an LLM-driven spoken dialogue system versus a human interlocutor. No significant overall differences were found, but the LLM condition showed slightly lower pronunciation and language use scores, slower speech rate, reduced lexical diversity yet higher syntactic complexity, and more initiative in discussion management. The findings suggest LLM-driven SDSs can serve as usable interlocutors but require refinement to better capture interactive listening.
Original abstract
The use of large language models (LLMs) in language assessment, particularly in spoken dialogue systems (SDSs) for assessing speaking, remains at an early stage. This study explored the use of an LLM-driven SDS to assess second language speaking ability. Using a within-participant design, we compared the paired discussion performance of 30 participants interacting with a self-built LLM-driven SDS (E-Talk) versus a human interlocutor, focusing on the interlocutor effect on test scores and fine-grained linguistic and interactional features. Results did not yield substantive differences in oral performance across the two conditions, although the interactions with the LLM-driven SDS displayed slightly lower scores for pronunciation and language use, slower speech rate, and reduced lexical diversity alongside increased syntactic complexity, and greater initiative in introducing new ideas, prompting responses, and guiding discussions toward negotiation, coupled with weaker interactive listening. These findings suggest that LLM-driven SDSs can serve as a potential, usable interlocutor in dialogic speaking assessments, eliciting key aspects of speaking ability. That said, further refinement is needed to better capture interactive listening and collaborative meaning construction, highlighting the importance of interpreting SDS-based performance relative to its specific interactional affordances.
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of ChatGPT-powered chatbot interactions in a low-stakes formative speaking assessment. Comparing a 290,000-word chatbot corpus with the British National Corpus 2014 revealed that chatbot output resembled written language more than spoken, with higher lexical density and fewer spoken features like stance markers. The framework is designed to be adaptable to future AI conversational agents.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.