language-assessment
Filtering by topic language-assessment(7)Clear all filters
- PaperIELTS Partnership Research Reports
An investigation of the language assessment interests and needs of professional registration bodies in the UK: An unconsidered perspective?
This study explores the language assessment interests and needs of professional registration bodies in the UK, highlighting a previously overlooked perspective in the field.
- PaperIELTS Partnership Research Reports
Comparing New TOEFL 2026 with Former TOEFL 2023 and IELTS
This paper compares the upcoming TOEFL 2026 test with its 2023 predecessor and the IELTS, highlighting key differences in structure and scoring.
- PaperOpenAlex — TESOL researchForthcoming · 1 Dec 2026
Future of AI in English Language Education: Trends and Predictions
Prof. Jagadeesh Nerlekar, K Munianjinappa
This paper synthesizes current AI developments in English language education, identifying key trends such as personalized feedback, adaptive pathways, automated formative assessment, realistic speaking/listening practice via conversational agents, and AI-assisted material creation. It predicts that hybrid systems combining large language models with pedagogical scaffolding and teacher mediation will have the most near-term impact. The paper also addresses challenges including bias, privacy, over-reliance on automation, and the need for teacher training, offering recommendations for educators and policymakers.
Original abstract
Artificial intelligence (AI) is transforming English language education (ELE) by enabling personalized learning, automated assessment, adaptive content generation, and immersive practice environments. This paper synthesizes current developments, identifies emergent trends, and offers evidence-informed predictions about how AI will shape classroom practice, curriculum design, assessment, teacher roles, and policy over the next decade. Drawing on interdisciplinary literature from computer-assisted language learning (CALL), intelligent tutoring systems (ITS), natural language processing (NLP), and educational policy, the paper argues that the most significant near-term impact will stem from hybrid systems that combine large language models (LLMs) with pedagogically informed scaffolding and teacher mediation. Key trends discussed include (1) ubiquitous personalized feedback and adaptive pathways; (2) automated, formative assessment with rich analytics; (3) realistic speaking/listening practice via multimodal conversational agents and immersive virtual environments; (4) AI-assisted material creation and differentiation for diverse learner needs; and (5) data-driven teacher support and professional development. Predictions address likely improvements in scalability and access, as well as persistent challenges: bias and fairness in language models, privacy and data governance, over-reliance on automated feedback, and the need for robust teacher training and curricular alignment. The paper concludes with practical recommendations for educators, institutions, and policymakers to harness AI’s affordances while safeguarding equity, transparency, and pedagogical quality. These include adopting hybrid human–AI workflows, emphasizing explainability and interpretability in tools, developing clear data-ethics policies, investing in teacher capacity building, and prioritizing research-practice partnerships. The analysis aims to be actionable for practitioners and decision-makers planning for an AI-augmented future of English language learning. Keywords: artificial intelligence, English language education, adaptive learning, large language models, assessment, teacher role, ethics
- PaperApplied Linguistics24 Jul 2026
INTRODUCING SECOND LANGUAGE ASSESSMENT
Lin Shi, Lianzhen He
An introduction to second language assessment is provided, covering key concepts and approaches.
- PaperarXiv — AI in Education (cs.CY)9 Jul 2026
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight, Danielle Carvalho et al.
L2-Bench, an open-source benchmark of 1000+ task-response pairs, measures LLM capabilities in applying learning experience design principles for second language education and assessment. Validated by over 200 practitioners, it evaluates 12 competencies across 31 subcompetencies using a rubric-based methodology. Results show Claude Opus 4.7 achieves the highest overall score (85.5%), but performance drops substantially on harder tasks (69.9% to 73.4%).
Original abstract
Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity: 4.42/5.00, criteria adequacy: 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks (69.9% to 73.4%). L2-Bench provides education stakeholders better methods to make more informed decisions about real-world AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.
- PaperDOAJ — Language assessment1 Jul 2026
Multilevel auditory and cognitive processing in post-stroke aphasia: associations with language performance using ABR, LLR, and P300: a cross-sectional observational study
Agit Şimşek, Nihal Sümeyye Ulutaş, Feyza Deniz Saman
The study compared auditory evoked potentials (ABR, LLR, P300) between post-stroke aphasia patients and healthy controls, finding significantly prolonged latencies across subcortical, cortical, and cognitive levels in aphasia. These latency delays indicate generalized neural slowing rather than reduced recruitment, suggesting that electrophysiological measures may complement behavioral language assessments in aphasia evaluation.
Original abstract
ABSTRACT BACKGROUND: Aphasia following stroke is primarily characterized by language impairment; however, accumulating evidence suggests that deficits in auditory and cognitive processing may also contribute to impaired language function. Auditory evoked potentials provide objective markers of neural processing across subcortical, cortical, and cognitive levels and may help clarify the neurophysiological mechanisms underlying post-stroke aphasia. OBJECTIVES: To investigate multilevel auditory processing and its relationship with language performance in individuals with post-stroke aphasia using Auditory Brainstem Responses (ABR), Late Latency Responses (LLR), and P300 potentials. DESIGN AND SETTING: Cross-sectional observational study conducted at a university hospital in Türkiye. METHODS: Twenty-nine individuals with post-stroke aphasia and 33 age-matched healthy controls were included. Language performance was assessed using the Aphasia Language Assessment Test and a standardized naming task. Auditory processing was evaluated using ABR, LLR (P1, N1, P2, N2), and P300 potentials recorded according to standard electrophysiological protocols. Group comparisons were performed for latency and amplitude measures, and correlation analyses were conducted to examine associations between electrophysiological parameters and language performance. RESULTS: Compared with healthy controls, individuals with aphasia demonstrated significantly prolonged latencies in ABR waves and interpeak intervals, as well as in all LLR and P300 components (p < 0.05). No significant group differences were observed in amplitude measures. Naming performance was significantly lower in the aphasia group. Although correlations between language performance and electrophysiological measures did not reach statistical significance, moderate negative trends were observed between naming scores and N2 and P300 latencies. CONCLUSION: Post-stroke aphasia is associated with delayed auditory processing across subcortical, cortical, and cognitive levels, reflecting generalized slowing rather than reduced neural recruitment. Latency-based auditory evoked potential measures may complement behavioral language assessments and support a multilevel auditory–cognitive framework for aphasia evaluation and rehabilitation.
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of chatbot simulated spoken interaction compared to natural human speech. Using a corpus of ChatGPT output and the British National Corpus, they found that chatbot production more closely resembles written than spoken language, with higher lexical density and fewer spoken features like stance markers. The framework is intended for use in developing and validating AI-powered conversational agents for language assessment.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.