test-validity
Filtering by topic test-validity(2)Clear all filters
- PaperLanguage Testing24 Jul 2026
Investigating the Real-World Relevance of an Academic English Speaking Test: Extrapolating Subjective Evaluations and Linguistic Performance Characteristics
Daniel R. Isbell, Dustin Crowther, Jieun Kim, Yoonseo Kim
The study examined correlations between TOEFL Essentials speaking scores and linguistic characteristics with academic speaking tasks in lab and course settings. Strong correlations were found, particularly for fluency and accuracy, supporting the extrapolation of test performances to academic contexts.
Original abstract
To support use of tests in academic contexts, it is critical to demonstrate that test scores and test performances are associated with performance in academic settings—an inferential link referred to as extrapolation in argument-based validation frameworks. TOEFL Essentials is a newer test designed to measure both general and academic English and is intended for use in higher education. The TOEFL Essentials speaking section consists of Virtual Interview, Read Aloud, and Listen & Repeat tasks, the latter two of which elicit highly constrained responses that may be less reflective of academic speaking tasks. In this study, we examined correlations of scores and linguistic characteristics across TOEFL Essentials speaking performances and (a) lab-based academic tasks (graph description, lecture response) for 149 students and (b) an authentic course-based speaking task for 65 students. Strong correlations (.65 < r < .80) were found between TOEFL Essentials speaking scores and evaluations of academic speaking. Among linguistic characteristics, fluency and accuracy variables demonstrated the largest and most consistent correlations across test and non-test tasks. Findings provide evidence relevant to the extrapolation of TOEFL Essentials speaking performances, which are based in part on highly constrained tasks, to academic settings and help inform decisions about test use.
- PaperLanguage Testing19 Jul 2026
Outcomes of One-Skill Retakes in a Four-Skills Proficiency Test: Evidence From Large-Scale Test Data
Hye-won Lee, Emma Bruce, Jan Langeslag, Reza Tasviri
Analysis of over 20,000 IELTS One Skill Retake (OSR) test takers from 2022 to Spring 2024 shows that retake component scores are on average higher than original scores, and overall band score changes are similar to those of short-interval full-test repeaters. Survey responses from 578 test takers identify insufficient preparation, stress, anxiety, and fatigue as common reasons for initial underperformance, providing empirical evidence for interpreting one-skill retake scores and discussing validity and equity of retake policies.
Original abstract
A test taker may underperform for reasons not fully attributable to language proficiency, including psychological or contextual influences such as anxiety or illness. IELTS One Skill Retake (OSR) was launched in 2022, allowing test takers to retake, within 60 days, a single component in which their initial performance may have been affected by extenuating circumstances. This study examines outcomes associated with OSR by analysing test-taking patterns and score changes among over 20,000 OSR test takers from its launch through Spring 2024. It also reports survey findings from 578 OSR test takers on their experiences, including whether they achieved target scores and their perceived reasons for not achieving the desired score on the original full test. Across skills, average OSR component scores were higher than the corresponding scores on the original full test, and overall-band changes among OSR test takers were comparable to those observed among short-interval full-test repeaters (⩽60 days). Survey responses commonly cited factors such as insufficient preparation, stress and anxiety, and fatigue and lack of focus as perceived contributors to underperformance on the initial test. The findings contribute empirical evidence relevant to the interpretation and use of one-skill retake scores and to ongoing discussions of the validity and equity implications of retake policies in large-scale language testing.