language-assessment
Filtering by topic language-assessment(12)Clear all filters
- PaperIELTS Partnership Research Reports
An investigation of the language assessment interests and needs of professional registration bodies in the UK: An unconsidered perspective?
This study explores the language assessment interests and needs of professional registration bodies in the UK, highlighting a previously overlooked perspective in the field.
- PaperIELTS Partnership Research Reports
Comparing New TOEFL 2026 with Former TOEFL 2023 and IELTS
This paper compares the upcoming TOEFL 2026 test with its 2023 predecessor and the IELTS, highlighting key differences in structure and scoring.
- PaperOpenAlex — TESOL researchForthcoming · 1 Dec 2026
Future of AI in English Language Education: Trends and Predictions
Prof. Jagadeesh Nerlekar, K Munianjinappa
This paper synthesizes current AI developments in English language education, identifying key trends such as personalized feedback, adaptive pathways, automated formative assessment, realistic speaking/listening practice via conversational agents, and AI-assisted material creation. It predicts that hybrid systems combining large language models with pedagogical scaffolding and teacher mediation will have the most near-term impact. The paper also addresses challenges including bias, privacy, over-reliance on automation, and the need for teacher training, offering recommendations for educators and policymakers.
Original abstract
Artificial intelligence (AI) is transforming English language education (ELE) by enabling personalized learning, automated assessment, adaptive content generation, and immersive practice environments. This paper synthesizes current developments, identifies emergent trends, and offers evidence-informed predictions about how AI will shape classroom practice, curriculum design, assessment, teacher roles, and policy over the next decade. Drawing on interdisciplinary literature from computer-assisted language learning (CALL), intelligent tutoring systems (ITS), natural language processing (NLP), and educational policy, the paper argues that the most significant near-term impact will stem from hybrid systems that combine large language models (LLMs) with pedagogically informed scaffolding and teacher mediation. Key trends discussed include (1) ubiquitous personalized feedback and adaptive pathways; (2) automated, formative assessment with rich analytics; (3) realistic speaking/listening practice via multimodal conversational agents and immersive virtual environments; (4) AI-assisted material creation and differentiation for diverse learner needs; and (5) data-driven teacher support and professional development. Predictions address likely improvements in scalability and access, as well as persistent challenges: bias and fairness in language models, privacy and data governance, over-reliance on automated feedback, and the need for robust teacher training and curricular alignment. The paper concludes with practical recommendations for educators, institutions, and policymakers to harness AI’s affordances while safeguarding equity, transparency, and pedagogical quality. These include adopting hybrid human–AI workflows, emphasizing explainability and interpretability in tools, developing clear data-ethics policies, investing in teacher capacity building, and prioritizing research-practice partnerships. The analysis aims to be actionable for practitioners and decision-makers planning for an AI-augmented future of English language learning. Keywords: artificial intelligence, English language education, adaptive learning, large language models, assessment, teacher role, ethics
- PaperApplied Linguistics24 Jul 2026
INTRODUCING SECOND LANGUAGE ASSESSMENT
Lin Shi, Lianzhen He
An introduction to second language assessment is provided, covering key concepts and approaches.
- PaperarXiv — AI in Education (cs.CY)9 Jul 2026
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight, Danielle Carvalho et al.
L2-Bench, an open-source benchmark of 1000+ task-response pairs, measures LLM capabilities in applying learning experience design principles for second language education and assessment. Validated by over 200 practitioners, it evaluates 12 competencies across 31 subcompetencies using a rubric-based methodology. Results show Claude Opus 4.7 achieves the highest overall score (85.5%), but performance drops substantially on harder tasks (69.9% to 73.4%).
Original abstract
Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity: 4.42/5.00, criteria adequacy: 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks (69.9% to 73.4%). L2-Bench provides education stakeholders better methods to make more informed decisions about real-world AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.
- PaperDOAJ — Language assessment1 Jul 2026
Multilevel auditory and cognitive processing in post-stroke aphasia: associations with language performance using ABR, LLR, and P300: a cross-sectional observational study
Agit Şimşek, Nihal Sümeyye Ulutaş, Feyza Deniz Saman
The study compared auditory evoked potentials (ABR, LLR, P300) between post-stroke aphasia patients and healthy controls, finding significantly prolonged latencies across subcortical, cortical, and cognitive levels in aphasia. These latency delays indicate generalized neural slowing rather than reduced recruitment, suggesting that electrophysiological measures may complement behavioral language assessments in aphasia evaluation.
Original abstract
ABSTRACT BACKGROUND: Aphasia following stroke is primarily characterized by language impairment; however, accumulating evidence suggests that deficits in auditory and cognitive processing may also contribute to impaired language function. Auditory evoked potentials provide objective markers of neural processing across subcortical, cortical, and cognitive levels and may help clarify the neurophysiological mechanisms underlying post-stroke aphasia. OBJECTIVES: To investigate multilevel auditory processing and its relationship with language performance in individuals with post-stroke aphasia using Auditory Brainstem Responses (ABR), Late Latency Responses (LLR), and P300 potentials. DESIGN AND SETTING: Cross-sectional observational study conducted at a university hospital in Türkiye. METHODS: Twenty-nine individuals with post-stroke aphasia and 33 age-matched healthy controls were included. Language performance was assessed using the Aphasia Language Assessment Test and a standardized naming task. Auditory processing was evaluated using ABR, LLR (P1, N1, P2, N2), and P300 potentials recorded according to standard electrophysiological protocols. Group comparisons were performed for latency and amplitude measures, and correlation analyses were conducted to examine associations between electrophysiological parameters and language performance. RESULTS: Compared with healthy controls, individuals with aphasia demonstrated significantly prolonged latencies in ABR waves and interpeak intervals, as well as in all LLR and P300 components (p < 0.05). No significant group differences were observed in amplitude measures. Naming performance was significantly lower in the aphasia group. Although correlations between language performance and electrophysiological measures did not reach statistical significance, moderate negative trends were observed between naming scores and N2 and P300 latencies. CONCLUSION: Post-stroke aphasia is associated with delayed auditory processing across subcortical, cortical, and cognitive levels, reflecting generalized slowing rather than reduced neural recruitment. Latency-based auditory evoked potential measures may complement behavioral language assessments and support a multilevel auditory–cognitive framework for aphasia evaluation and rehabilitation.
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of chatbot simulated spoken interaction compared to natural human speech. Using a corpus of ChatGPT output and the British National Corpus, they found that chatbot production more closely resembles written than spoken language, with higher lexical density and fewer spoken features like stance markers. The framework is intended for use in developing and validating AI-powered conversational agents for language assessment.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.
- PaperLanguage Testing19 Jun 2026
Book Review: Assessment of Plurilingual Competence and Plurilingual Learners in Educational Settings: Educative Issues and Empirical Approaches Melo-PfeiferS.OllivierC. (Eds.), Assessment of Plurilingual Competence and Plurilingual Learners in Educational Settings: Educative Issues and Empirical Approaches. Routledge, 2024. 272 pp. ISBN 9781032011097 (hbk) US$200.00 ISBN 9781032011103 (sbk) US$55.99 ISBN 9781003177197 (ebk) US$41.99
Gordon Blaine West
A review of a 2024 edited volume on assessing plurilingual competence in educational settings, covering both theoretical and empirical approaches.
- PaperETS Research Report Series16 Jun 2026
Mapping TOEIC® Link™ Scores to the Common European Framework of Reference for Languages
Kathryn Hille, Jonathan Schmidgall, Saerhim Oh, Renka Ohta
This study maps TOEIC® Link™ assessment scores to the Common European Framework of Reference for Languages using expert judgment. The Bookmark method was applied for multiple-choice listening and reading items, while the Expected Task Score method was used for constructed-response speaking and writing items. Panelists reported high agreement and satisfaction with the cut scores, with minor post-study adjustments for consistency across ETS assessments.
Original abstract
This study maps TOEIC® Link™ assessment scores to levels of the Common European Framework of Reference for Languages based on the judgments of a panel of experts. The Bookmark method was used to map the TOEIC Link assessments containing multiple-choice items (listening and reading), whereas the Expected Task Score method was used to map the assessments containing constructed-response items (speaking and writing). The panelists reported high levels of agreement, comfort with the panelist-recommended cut scores, and satisfaction with the standard setting process. A few small poststudy adjustments were made in the interest of congruence among ETS English proficiency assessments with common items or overlapping task types. Suggested citation: Hille, K., Schmidgall, J., Oh, S., & Ohta, R. (2026). Mapping TOEIC® Link™ scores to the Common European Framework of Reference for Languages (Research memorandum No. RM-26-04). ETS. https://doi.org/10.64634/4amdns26
- PaperAssessing Writing2 Jun 2026
Reassessing automated essay scoring with large language models: Evidence from API and GUI interfaces
Jieun Kim, Daniel Holden
This study re-examines the effectiveness of automated essay scoring using large language models by comparing results from API and GUI interfaces, providing evidence on the impact of interface design on scoring accuracy.
- PaperDOAJ — Language assessment1 Jun 2026
Glocalizing translation assessment: insights from the College English Test
Lin Zhang, Yan Jin
Using the College English Test (CET) in China as a case study, this article demonstrates how translation assessment can be glocalized – aligning test content with local socio-cultural contexts while maintaining global quality standards. The study analyzes the evolution of the CET translation component, showing how Chinese themes and discourse structures are integrated, and provides empirical validity evidence for the revised tasks.
Original abstract
Abstract This article examines the glocalization of translation assessment through a case study of the translation component of the College English Test (CET) in China. Specifically, the study investigates how the CET, through the reform of its translation component, sought to align its test content with the local context while maintaining global standards of quality. Our case study reveals that the content and format of the CET translation assessment have evolved in response to the changing societal and educational needs. Through an in-depth analysis of task design, the study illustrates how Chinese socio-cultural themes and discourse structures are incorporated to enhance the local relevance of the test. At the same time, the quality assurance measures implemented by the test developer are described in detail, and empirical evidence is presented to support the validity of the revised task. Our study illuminates how large-scale language assessments can integrate national education priorities with international standards of quality and fairness, thereby offering insights for language assessment professionals seeking to reconcile local responsiveness with global quality expectations.
- PaperETS Research Report Series29 May 2026
Alignment of the TOEIC Bridge® Tests With the Educational Functioning Level Descriptors for English as a Second Language
Mikyung Kim Wolf, Renka Ohta
This alignment study evaluated how well the TOEIC Bridge tests correspond to the NRS educational functioning level descriptors for ESL. Results showed strong content alignment across test forms, providing validity evidence for using these assessments in adult education programs for NRS accountability. The report offers methodology and recommendations for effective use.
Original abstract
This report presents the findings of an alignment study evaluating the extent to which the TOEIC Bridge® tests reflect the language knowledge, skills, and abilities described in the National Reporting System (NRS) educational functioning level descriptors for English as a second language (ESL). The study was intended to provide critical validity evidence supporting the use of these assessments in adult education programs serving ESL learners, particularly for purposes of NRS accountability. Results indicate strong evidence of content alignment across all test forms examined. The report details the research-based methodology and offers practical recommendations for the effective use of the TOEIC Bridge tests in the NRS. Suggested citation: Wolf, M. K., & Ohta, R. (2026). Alignment of the TOEIC Bridge® tests with the educational functioning level descriptors for English as a second language (Research Memorandum No. RM-26-03). ETS. https://doi.org/10.64634/1n934x31