language-testing
Filtering by topic language-testing(5)Clear all filters
- PaperIELTS Partnership Research Reports
A concordance study of CELPIP General and IELTS General Training
A concordance analysis of the CELPIP General and IELTS General Training tests was conducted to evaluate score alignment between the two assessments.
- PaperIELTS Partnership Research Reports
Safeguarding equity, access and inclusion in IELTS
This paper examines strategies to ensure fairness and inclusivity in the IELTS test, addressing potential biases and barriers for diverse test-takers.
- PaperIELTS Partnership Research Reports
Aligning scores of language proficiency tests: A score concordance study between IELTS Academic and TOEFL iBT
This study establishes score concordance tables between IELTS Academic and TOEFL iBT, enabling direct comparison of proficiency levels across the two tests.
- PaperLanguage Testing19 Jul 2026
Outcomes of One-Skill Retakes in a Four-Skills Proficiency Test: Evidence From Large-Scale Test Data
Hye-won Lee, Emma Bruce, Jan Langeslag, Reza Tasviri
The study analyzes over 20,000 test takers using IELTS One Skill Retake (OSR) and finds that average retake scores were higher than original scores, with overall band changes similar to short-interval full-test repeaters. Survey data from 578 test takers indicates factors like insufficient preparation, stress, and fatigue contributed to initial underperformance. The findings provide empirical evidence for interpreting one-skill retake scores and inform discussions on validity and equity in large-scale language testing.
Original abstract
A test taker may underperform for reasons not fully attributable to language proficiency, including psychological or contextual influences such as anxiety or illness. IELTS One Skill Retake (OSR) was launched in 2022, allowing test takers to retake, within 60 days, a single component in which their initial performance may have been affected by extenuating circumstances. This study examines outcomes associated with OSR by analysing test-taking patterns and score changes among over 20,000 OSR test takers from its launch through Spring 2024. It also reports survey findings from 578 OSR test takers on their experiences, including whether they achieved target scores and their perceived reasons for not achieving the desired score on the original full test. Across skills, average OSR component scores were higher than the corresponding scores on the original full test, and overall-band changes among OSR test takers were comparable to those observed among short-interval full-test repeaters (⩽60 days). Survey responses commonly cited factors such as insufficient preparation, stress and anxiety, and fatigue and lack of focus as perceived contributors to underperformance on the initial test. The findings contribute empirical evidence relevant to the interpretation and use of one-skill retake scores and to ongoing discussions of the validity and equity implications of retake policies in large-scale language testing.
- PaperLanguage Testing19 Jul 2026
Can an AI Agent Replace Human Examiners in High-Stakes Interactive Speaking Tests? A Debate
Jing Xu, Lynda Taylor, Xiaoming Xi, Yasin Karatay et al.
This viewpoint article presents arguments for and against replacing human examiners with AI agents in high-stakes interactive speaking tests, drawing on a debate at the 2025 LTRC. It examines construct theory, practicality, washback, fairness, and ethical uses of AI in language assessment.
Original abstract
Generative artificial intelligence (GenAI) is advancing at a remarkable speed, and its promise in transforming current practice in language assessment has been articulated by applied linguistics researchers. An emerging application of GenAI is to integrate the technology into Spoken Dialogue Systems (SDSs) to simulate human interlocutors for the purpose of speaking practice or assessment. Despite rapid technological advances, the issue of whether an AI agent can replace a human examiner in one-on-one, high-stakes interactive speaking tests remains contentious. Building on a lively academic debate on this topic at the 2025 Language Testing Research Colloquium (LTRC) in Bangkok, this Viewpoint presents arguments both for and against this proposition in terms of construct theory, practicality, washback, fairness, ethical uses of AI, and so forth.