speaking-assessment
Filtering by topic speaking-assessment(6)Clear all filters
- PaperLanguage Testing24 Jul 2026
Investigating the Real-World Relevance of an Academic English Speaking Test: Extrapolating Subjective Evaluations and Linguistic Performance Characteristics
Daniel R. Isbell, Dustin Crowther, Jieun Kim, Yoonseo Kim
This study examined the real-world relevance of TOEFL Essentials speaking scores by correlating them with performance on lab-based and authentic academic speaking tasks among 149 and 65 students, respectively. Strong correlations were found, especially for fluency and accuracy measures, supporting the extrapolation of test scores to academic settings.
Original abstract
To support use of tests in academic contexts, it is critical to demonstrate that test scores and test performances are associated with performance in academic settings—an inferential link referred to as extrapolation in argument-based validation frameworks. TOEFL Essentials is a newer test designed to measure both general and academic English and is intended for use in higher education. The TOEFL Essentials speaking section consists of Virtual Interview, Read Aloud, and Listen & Repeat tasks, the latter two of which elicit highly constrained responses that may be less reflective of academic speaking tasks. In this study, we examined correlations of scores and linguistic characteristics across TOEFL Essentials speaking performances and (a) lab-based academic tasks (graph description, lecture response) for 149 students and (b) an authentic course-based speaking task for 65 students. Strong correlations (.65 < r < .80) were found between TOEFL Essentials speaking scores and evaluations of academic speaking. Among linguistic characteristics, fluency and accuracy variables demonstrated the largest and most consistent correlations across test and non-test tasks. Findings provide evidence relevant to the extrapolation of TOEFL Essentials speaking performances, which are based in part on highly constrained tasks, to academic settings and help inform decisions about test use.
- PaperRELC Journal21 Jul 2026
Enhancing English Public Speaking Self-Efficacy and Performance Through a CDIO-Based English Course
Yen-Hui Lu
Integrating the CDIO (Conceive-Design-Implement-Operate) framework into English public speaking instruction for EFL university students significantly improved their self-efficacy and performance in content development, organization, and delivery, but not language accuracy or fluency. The quasi-experimental study with 18-week intervention showed that project-driven tasks like innovation pitches fostered mastery experiences. Pedagogical implications include using CDIO's four-phase structure to scaffold learning and supplementing with language-focused interventions to address accuracy.
Original abstract
This study investigates the impact of integrating the Conceive–Design–Implement–Operate (CDIO) framework into English public speaking (EPS) instruction for English as a foreign language (EFL) university students in Taiwan. EPS has been widely recognized as a critical competence in academic and professional contexts, yet EFL learners often struggle with anxiety, limited practice opportunities and low self-efficacy. To address these challenges, the study implemented a quasi-experimental design with two Freshman English classes: a control group receiving traditional presentation training; and an experimental group engaged in a CDIO-based presentation training. Over an 18-week semester, students’ self-efficacy was measured using a validated EPS self-efficacy scale, while performance was evaluated with a rubric assessing language use, content, organization and delivery. Pre-test results showed no significant differences between groups, ensuring baseline equivalence. Post-test findings revealed that the experimental group demonstrated significantly higher self-efficacy, as well as superior performance in content development, organizational structure and delivery techniques. However, no significant improvements were found in language accuracy and fluency. These results suggest that CDIO-based instruction enhances students’ confidence and competence in public speaking through authentic, project-driven tasks. The findings yield three key pedagogical implications: (a) integrating authentic, product-based tasks such as sustainable innovation pitches can foster mastery experiences and boost learner motivation and self-efficacy; (b) using CDIO's four-phase structure enables instructors to scaffold content and organizational development through guided brainstorming, outlining and rehearsed delivery; and (c) to address limitations in grammar and vocabulary acquisition, supplementary language-focused interventions such as vocabulary mini-tasks and grammar workshops should be embedded within the CDIO cycle to balance fluency with accuracy. These strategies can optimize the effectiveness of CDIO-based EPS instruction in EFL contexts.
- PaperLanguage Testing8 Jul 2026
Assessing Interactional Competence Through Generative AI: Comparing Large Language Models as AI Interlocutors in the Paired Oral Discussion Test
Inyoung Na
This study compared GPT-4o and Claude 3.5 Sonnet as AI interlocutors in a paired oral discussion test to assess interactional competence (IC). Claude outperformed GPT-4o in eliciting IC features such as communication breakdown strategies and stance maintenance, and was perceived as more authentic by test takers. The findings underscore the importance of construct-driven evaluation when selecting large language models for language assessment contexts.
Original abstract
Interactional competence (IC) is essential for oral communication assessment, yet human partner variability can introduce construct-irrelevant variance in paired speaking tests. As an alternative to a test with a human interlocutor, this study describes the development of a large language model (LLM)-driven Spoken Dialogue System and compares GPT-4o to Claude 3.5 Sonnet to inform model selection for IC assessment. Twelve international students completed paired discussion tasks with both LLMs in counterbalanced order. System performance was evaluated through breakdown activation consistency, stance maintenance, and persona adherence. Test-taker performances were analyzed using interactional discourse analysis to identify IC features across three dimensions: topic management, interactional management, and interactive listening. Semi-structured interviews explored test takers’ perceptions of the AI partners. Results showed Claude outperformed GPT-4o in eliciting IC features, successfully activating communication breakdown strategies and maintaining oppositional stance, thereby creating more opportunities for test takers to demonstrate key IC abilities. Test takers perceived Claude as more authentic and natural, while GPT was perceived as more artificial. These findings demonstrate that different LLMs create distinct interactional conditions affecting both IC elicitation and test-taker perceptions. The findings highlight the need for construct-driven evaluation criteria when selecting LLMs for language-assessment contexts.
- PaperLanguage Testing6 Jul 2026
Human or Machine? Evaluating Second Language Speaking Performance in Paired Discussions with a Large Language Model-Driven Spoken Dialogue System vs. a Human Interlocutor
Shangchao Min, Zhuohan Hou, Yanxin Wang
A within-participant study compared 30 L2 learners' paired discussion performance with an LLM-driven spoken dialogue system versus a human interlocutor. No substantial score differences emerged, but the LLM condition yielded slightly lower pronunciation and language use scores, slower speech, less lexical diversity, greater syntactic complexity, and more proactive topic initiation with weaker interactive listening. The findings suggest LLM-driven SDSs can serve as viable interlocutors in dialogic speaking assessment, though refinement is needed to better capture interactive listening and collaborative meaning construction.
Original abstract
The use of large language models (LLMs) in language assessment, particularly in spoken dialogue systems (SDSs) for assessing speaking, remains at an early stage. This study explored the use of an LLM-driven SDS to assess second language speaking ability. Using a within-participant design, we compared the paired discussion performance of 30 participants interacting with a self-built LLM-driven SDS (E-Talk) versus a human interlocutor, focusing on the interlocutor effect on test scores and fine-grained linguistic and interactional features. Results did not yield substantive differences in oral performance across the two conditions, although the interactions with the LLM-driven SDS displayed slightly lower scores for pronunciation and language use, slower speech rate, and reduced lexical diversity alongside increased syntactic complexity, and greater initiative in introducing new ideas, prompting responses, and guiding discussions toward negotiation, coupled with weaker interactive listening. These findings suggest that LLM-driven SDSs can serve as a potential, usable interlocutor in dialogic speaking assessments, eliciting key aspects of speaking ability. That said, further refinement is needed to better capture interactive listening and collaborative meaning construction, highlighting the importance of interpreting SDS-based performance relative to its specific interactional affordances.
- PaperRELC Journal19 Jun 2026
Assessing pragmatic competence in ELF spoken interaction: Analytic rubrics for a university programme
Paul McBride
This study proposes a framework of analytic rubrics for assessing pragmatic competence in English as a Lingua Franca (ELF) spoken interaction. The framework focuses on three constructs: intelligibility, linguistic accommodation, and communication strategies, drawing on ELF research and corpus data from a university program in Tokyo. It challenges native-speaker norms and aims to make pragmatic competence more overt and assessable in language teaching programs.
Original abstract
English functions as a widespread communicative resource which is protean − highly adaptable in form − and shaped by contingencies of context. This situated variability constitutes clear evidence that pragmatic competence is central in effective language use. Understandings of English as an emergent practice therefore inform the pedagogical and administrative framework of an English as a Lingua Franca (ELF) programme at a private university in Tokyo. Within this framework, departures from native-speaker norms are interpreted as expressions of the diversity of global English use rather than as deficiencies. Despite the importance of pragmatic competence for successful ELF interaction, its role in assessment within the programme has remained largely implicit, a limitation which forms the basis of the rationale for the present study. Accordingly, this study proposes a framework of analytic rubrics aimed at making pragmatic competence in speaking more overt and assessable, expressed through attention to three interrelated constructs: intelligibility; linguistic accommodation; and communication strategies. Founded on ELF research and established principles of language assessment, the framework emphasizes clarity-oriented intelligibility work, adaptive linguistic accommodation and multistep meaning−negotiation strategies, reflecting salient pragmatic processes identified in ELF research. Each construct is exemplified by samples drawn from the English as a Lingua Franca in Japan Corpus, a resource developed by teacher-researchers in the programme. The framework therefore challenges evaluative paradigms which reward linguistic ‘correctness’ and risk marginalizing multilingual speakers whose context-specific language use facilitates successful transcultural communication regardless of conformity to native norms. Although not yet implemented, its construct-oriented design is intended to support engagement with ELF-related pedagogies and to facilitate critical language awareness. The study also proposes future directions, including rubrics for written ELF and learner-facing iterations for self-assessment, while offering a reference point for integrating ELF-aware speaking assessment into English language teaching programmes.
- PaperDOAJ — ELT & TESOL1 Jun 2022
Multimodal Learning Material in An English-Speaking Class in Kampung Baluwarti
Fitria Yuliani
This study examined the use of multimodal learning materials in an English speaking class for a tourism community in Kampung Baluwarti, Indonesia. Through a qualitative narrative design, it found that materials combining sound, images, motion, and text increased student engagement and comfort, leading to improved speaking skills.
Original abstract
This study mainly concerned about the use of Multimodal learning material in enhancing students’ English-speaking skill. The study was carried out in Kampung Baluwarti Surakarta as the researcher was one of the tutors involved in Community service team of ABA St. Pignatelli Surakarta during January to March 2022. The main objective of this study was to investigate the implementation of multimodal learning material used in English speaking class of Kampung Baluwarti tourism community. Furthermore, their perceptions towards the use of the learning material were depicted as well. Qualitative research with narrative design was used in this study. The researcher collected the data through a variety of data collection got from in-depth interviews with the students, observation, and documents. It was discovered that multimodal learning material could be presented in various mode such as sound, picture, motion and written text to accommodate students’ learning. It helped students to be more active in practicing the target language. Students’ involvement was seen more during the lesson. Moreover, students seemed to be excited and comfortable during the class. In other words, multimodal learning material applied in the teaching learning process makes a difference in students’ speaking skill achievement.