language-testing
Filtering by topic language-testing(21)Clear all filters
- PaperIELTS Partnership Research Reports
A concordance study of CELPIP General and IELTS General Training
A concordance analysis of the CELPIP General and IELTS General Training tests was conducted to evaluate score alignment between the two assessments.
- PaperIELTS Partnership Research Reports
Safeguarding equity, access and inclusion in IELTS
This paper examines strategies to ensure fairness and inclusivity in the IELTS test, addressing potential biases and barriers for diverse test-takers.
- PaperIELTS Partnership Research Reports
Aligning scores of language proficiency tests: A score concordance study between IELTS Academic and TOEFL iBT
This study establishes score concordance tables between IELTS Academic and TOEFL iBT, enabling direct comparison of proficiency levels across the two tests.
- PaperLanguage Testing19 Jul 2026
Outcomes of One-Skill Retakes in a Four-Skills Proficiency Test: Evidence From Large-Scale Test Data
Hye-won Lee, Emma Bruce, Jan Langeslag, Reza Tasviri
The study analyzes over 20,000 test takers using IELTS One Skill Retake (OSR) and finds that average retake scores were higher than original scores, with overall band changes similar to short-interval full-test repeaters. Survey data from 578 test takers indicates factors like insufficient preparation, stress, and fatigue contributed to initial underperformance. The findings provide empirical evidence for interpreting one-skill retake scores and inform discussions on validity and equity in large-scale language testing.
Original abstract
A test taker may underperform for reasons not fully attributable to language proficiency, including psychological or contextual influences such as anxiety or illness. IELTS One Skill Retake (OSR) was launched in 2022, allowing test takers to retake, within 60 days, a single component in which their initial performance may have been affected by extenuating circumstances. This study examines outcomes associated with OSR by analysing test-taking patterns and score changes among over 20,000 OSR test takers from its launch through Spring 2024. It also reports survey findings from 578 OSR test takers on their experiences, including whether they achieved target scores and their perceived reasons for not achieving the desired score on the original full test. Across skills, average OSR component scores were higher than the corresponding scores on the original full test, and overall-band changes among OSR test takers were comparable to those observed among short-interval full-test repeaters (⩽60 days). Survey responses commonly cited factors such as insufficient preparation, stress and anxiety, and fatigue and lack of focus as perceived contributors to underperformance on the initial test. The findings contribute empirical evidence relevant to the interpretation and use of one-skill retake scores and to ongoing discussions of the validity and equity implications of retake policies in large-scale language testing.
- PaperLanguage Testing19 Jul 2026
Can an AI Agent Replace Human Examiners in High-Stakes Interactive Speaking Tests? A Debate
Jing Xu, Lynda Taylor, Xiaoming Xi, Yasin Karatay et al.
This viewpoint article presents arguments for and against replacing human examiners with AI agents in high-stakes interactive speaking tests, drawing on a debate at the 2025 LTRC. It examines construct theory, practicality, washback, fairness, and ethical uses of AI in language assessment.
Original abstract
Generative artificial intelligence (GenAI) is advancing at a remarkable speed, and its promise in transforming current practice in language assessment has been articulated by applied linguistics researchers. An emerging application of GenAI is to integrate the technology into Spoken Dialogue Systems (SDSs) to simulate human interlocutors for the purpose of speaking practice or assessment. Despite rapid technological advances, the issue of whether an AI agent can replace a human examiner in one-on-one, high-stakes interactive speaking tests remains contentious. Building on a lively academic debate on this topic at the 2025 Language Testing Research Colloquium (LTRC) in Bangkok, this Viewpoint presents arguments both for and against this proposition in terms of construct theory, practicality, washback, fairness, ethical uses of AI, and so forth.
- PaperReCALL13 Jul 2026
The effectiveness of computerized dynamic assessment in improving L2 performance: A three-level meta-analysis
Qi Lu, Mengqi Chen, Lianrui Yang, Shaofeng Li et al.
A three-level meta-analysis of 35 studies (2000-2025) found that computerized dynamic assessment (C-DA) significantly improves L2 performance, with both cake-format (mediation embedded in tests) and sandwich-format (mediation between pretest and posttest) designs yielding large positive effects. The cake format showed larger effect sizes, and effectiveness was moderated by number of items, test content, and learners' first language.
Original abstract
The growing body of research on the effects of computerized dynamic assessment (C-DA) on second language (L2) learning underscores the need for a comprehensive research synthesis to identify future research directions and inform the application of C-DA in L2 educational contexts. This meta-analysis employed a three-level modeling approach to examine the effectiveness of C-DA in improving L2 learners’ performance. It synthesized 27 effect sizes from cake format designs, in which mediation is embedded within the test sequence, and 24 effect sizes from sandwich format designs, where mediation is delivered between a pretest and a posttest, across 35 studies published between 2000 and May 27, 2025. This study also investigated the key variables that moderate C-DA effectiveness. Findings reveal large, significant positive effects of both the cake and sandwich formats on L2 performance improvement (cake format: g = 2.120, p < .001; sandwich format: g = 1.676, p < .001), with the cake format tending to yield larger effect sizes. This may be because the cake format captures gains during mediation, whereas the sandwich format reflects post-mediation outcomes. Moderator analyses show that the number of items, test content, and learners’ first language affect C-DA effectiveness in promoting L2 performance. Drawing on the synthesized findings, this study contributes to theoretical, methodological, and technological understandings of C-DA and offers suggestions for future research in this domain.
- PaperLanguage Testing2 Jul 2026
Test Review: The Test of Proficiency in Korean
So-Young Lim, John Dylan Burton
This test review provides an overview of the TOPIK's history, purposes, design, and administration, and appraises its strengths and challenges. The lack of publicly available information on test construct, psychometric properties, and standard-setting procedures makes it difficult to fully evaluate score validity and reliability. Additional validation research is needed to examine score generalizability beyond academic domains.
Original abstract
As a nationally accredited test, the Test of Proficiency in Korean (TOPIK) assesses general Korean proficiency as a foreign language for various high-stakes purposes, such as university admissions, employment, and visa issuance. For this reason, the social impact of TOPIK on test takers is significant and cannot be underestimated. Despite the growing number of international test takers and stakeholders, there is limited validation research evaluating TOPIK’s various uses, as well as critical evaluation of the test itself. Thus, this test review provides an overview of the history, test purposes and use, design, and administration of TOPIK and offers an appraisal of its strengths and challenges. While the test scores are broadly utilized for their intended purposes and provide some evidence of language development across four skills in Korean, the lack of publicly available information on the test construct, psychometric properties, and standard-setting procedures makes it difficult to fully evaluate the validity and reliability of the scores. Furthermore, additional evidence and validation research are needed to examine the generalizability of TOPIK scores beyond academic domains, given the test’s diverse applications. Addressing these gaps is critical to meet the needs of various stakeholders and to strengthen the overall validity of the test.
- PaperDOAJ — Language assessment1 Jul 2026
The Design of Language Testing and Evaluation Materials for the English Department
Irra Wahidiyati, Windhariyati Dyah K, Ghaida Thifal
A needs analysis of English department students and lecturers informed the design of a 10-chapter textbook on language testing and evaluation. The material covers assessment concepts, principles, and specific techniques for listening, speaking, reading, writing, grammar, and vocabulary. The resulting resource aims to improve Language Assessment Literacy by integrating theory, test construction, and rubric development.
Original abstract
Background - The students of Tadris Bahasa Inggris need explicit material about assessments and tests. They need the material to develop listening and speaking skills. Next, the TBI students must master the creation of writing and reading test items. In addition to creating writing and reading test items. Urgency of Research – The needs analysis shows that students’ needs regarding the Language Testing and Evaluation subjects in the English Department of UIN Prof. K. H. Saifuddin Zuhri Purwokerto. They need information on how to create assessment rubrics for them. The learning materials used are compiled to meet the students' needs. Research Objectives – This study conducts a needs analysis of students' needs regarding the Language Testing and Evaluation subject and designs new material for it. Research Method - A mixed method was used to answer the research questions. The researcher analyzed the institution's problems and assessed whether students' needs aligned with the syllabus or lesson plan. The information sources were 100 students and 2 lecturers. Research Findings – The draft of the material consists of 7 chapters. Chapter 1 discusses assessment concepts and issues. Chapter 2 discusses the principles of language assessment. Chapter 3 discusses the design of classroom language tests and standardized testing. Chapter 5 discussed the assessment of listening, including intensive, responsive, selective, and extensive listening. Chapter 6 discusses assessing speaking, including imitative, intensive, responsive, interactive, and extensive speaking. Chapter 7 discusses the assessment of reading, including perceptive, selective, interactive, and extensive reading. Chapter 8 discusses assessing writing, including imitative, intensive, responsive, and extensive. Chapter 9 discusses the assessment of grammar and vocabulary. The last chapter discusses the grading and evaluation process. Research Conclusion & Novelty - The materials integrate assessment theory, test construction, and rubric development for all language skills into a single contextualized resource designed to promote Language Assessment Literacy.
- PaperLanguage Teaching1 Jul 2026
LTA volume 59 issue 3 Cover and Back matter
This item is the cover and back matter of Language Testing journal volume 59 issue 3, containing non-article administrative content.
- PaperDOAJ — Language assessment1 Jul 2026
A LINEAR LOGISTIC TEST MODEL (LLTM) APPLICATION IN FOREIGN LANGUAGE TESTING
Jose Fabián Elizondo-González, Peyman Jahanbin
This study applies the Linear Logistic Test Model (LLTM) to a reading comprehension subtest, finding that a Q-matrix of cognitive predictors (e.g., inferences, subtask complexity) explains 77% of item difficulty variance. The model enhances interpretability of item difficulty and identifies features for test refinement.
Original abstract
This study applies the Linear Logistic Test Model (LLTM) to the reading comprehension subtest of an English certification exam developed by the Foreign Language Assessment Program (PELEx, for its acronym in Spanish). A Q-matrix operationalized cognitive and linguistic predictors, such as paraphrasing and inferential reasoning, to explain item difficulty. Delta-squared results showed that the Q-matrix accounted for 77% of the variance in item difficulty, with “Inferences” and “Subtask Complexity” being key contributors. While overlaps in item difficulty coefficients reflected the nested nature of the Common European Framework of Reference for Languages (CEFR) levels, this progression aligns with the framework’s principles. The findings show that LLTM can make item difficulty more interpretable in reading assessment, while also helping to identify which item features merit further refinement in future test development.
- PaperLanguage Testing25 Jun 2026
Reflections on the Practical Implementation of Knoch and Fan’s (2024) Good Practice Principles for Score Concordance Studies
Spiros Papageorgiou, Tony Clark
Reviews the practical implementation of Knoch and Fan's (2024) good practice principles for score concordance studies, drawing on experience from a large-scale concordance study comparing IELTS Academic and TOEFL iBT scores. Emphasizes methodological rigor, transparency, construct comparability, and limitations of concordance tables for admissions decisions. Offers recommendations for fair score requirements regardless of test choice.
Original abstract
When different tests are used for the same purpose, score requirements should be comparable so that examinees cannot obtain an unfair advantage simply because of the test they chose. Drawing on our experience conducting a large-scale concordance study to allow for an empirical comparison of IELTS Academic and TOEFL iBT test scores, we review Knoch and Fan’s evaluative framework, explore methodological best practices and challenges, and offer future directions for score concordance research. We emphasize the importance of methodological rigor in collecting test-taker score data, transparency in analyzing such data to build score concordance tables, and a reasonable degree of construct comparability as a prerequisite for conducting a score concordance study, while also highlighting the limitations of concordance tables as standalone tools for admissions decisions. We note that some aspects of Knoch and Fan’s good practice principles are more straightforward to implement in practice than others. The good practice principles could be updated or adjusted after real-world application, which we describe with a view to furthering best practice in concordance research. We conclude this viewpoint with recommendations for decision-making that are based on fair score requirements, irrespective of which test the examinees chose.
- PaperLanguage Testing25 Jun 2026
Concordance or Discordance? The Broader Context of English Tests for Australian Immigration
John Read
The paper examines the alignment or misalignment of English language tests used for Australian immigration within their broader socio-political context.
- PaperLanguage Testing10 Jun 2026
Beyond Traditional Differential Item Functioning Detection: A Rasch Tree Approach to Evaluating Item Fairness in a Large-Scale German Reading Comprehension Test
Farshad Effatpanah, Olga Kunina-Habenicht, Katharina Antonia Michiko Tremmel, Philipp Sonnleitner
This study applies a recursive partitioning Rasch tree model to detect differential item functioning (DIF) in a large-scale German reading comprehension test for fifth graders, examining effects of gender, socioeconomic status, immigration status, personality, and need for cognition. Unlike traditional methods, the Rasch tree does not require pre-specified groups and can handle continuous covariates. Analysis of 4,252 students revealed nine non-predefined nodes, with eleven items showing moderate to large DIF, indicating that combinations of covariates impact test performance.
Original abstract
A key psychometric phenomenon in educational testing is differential item functioning (DIF), which evaluates whether specific test items function differently across subgroups of examinees who have the same level of the underlying (latent) ability. DIF happens when examinees with the same latent trait have different probabilities of correctly responding to a test item, influenced by their subgroup membership. This study aims to illustrate the application of the recursive partitioning Rasch tree model, shortly called the Rasch tree, to examine DIF in a large-scale German reading comprehension test for fifth-grade elementary students across gender, socioeconomic status, immigration status, personality disposition, and the need for cognition. Unlike the conventional DIF detection methods, the Rasch tree does not require a pre-specification of groups for exploring DIF, and continuous covariates can be easily included in the analysis. To investigate DIF of the test, item responses of 4,252 students were analyzed. The Rasch tree analysis generated nine non-predefined nodes, with slightly different patterns of item difficulties. Eleven items were flagged as exhibiting moderate and large DIF in the four splitting nodes. No splits were produced based on the need for cognition. The results indicated that the combination of the covariates impacted students’ test performance.
- PaperAssessing Writing2 Jun 2026
Examining the relevance of three TOEFL Essentials writing tasks to the accounting profession: The role of domain experts
Ute Knoch, Jason Fan, Michael Davey, Sally O’Hagan et al.
Domain experts (accountants) evaluated three TOEFL Essentials writing tasks for relevance to accounting. The Build a Sentence task was seen as least relevant, while Write an Email was most relevant, but experts focused on a narrow range of features and sometimes misinterpreted task demands.
Original abstract
Large-scale English language proficiency tests are increasingly used to make decisions about professional registration despite not originally being developed to make predictions about language use in the workplace. Domain experts can play a valuable role as informants in establishing the relevance of test tasks to a specific TLU domain. However, this practice has seldom been critically examined. In particular, no studies to date have examined the nature of the interview questions used when engaging with domain experts in this type of research. The current study was designed with two aims: to (1) explore the relevance of three writing tasks (i.e., Build a Sentence, Write for an Academic Discussion, Write an Email) from the TOEFL Essentials test to the accounting profession and (2) evaluate the judgements of domain expert participants. Twenty accountants from non-English speaking backgrounds as well as three accounting educators were interviewed for the study, drawing on a methodology with broad, general questions. The data was analysed qualitatively to identify (a) to what extent the participants considered the three tasks relevant and (b) what task features they attended to when commenting on the relevance of the tasks. The findings showed that the participants generally found the Build a Sentence the least relevant of the three writing tasks, and the Write an Email task the most relevant. When reviewing the task features the participants judged as relevant to writing demands in their workplace, it was shown they focused on a small/narrow range of features. They also engaged in ‘misinterpretations’, comparing aspects of test and workplace tasks that did not align (e.g., comparing a writing task to events in a spoken meeting). The findings are discussed in terms of domain expert involvement in validation research.
- PaperLanguage Testing21 May 2026
Book Review: Language Testing and Assessment: From Theory to Practice PhakitiA., Language Testing and Assessment: From Theory to Practice. Bloomsbury Academic, 2025. 312 pp. ISBN 978 1 47429 012 8 (paperback), £22.49 (paperback), £67.50 (hardback), £22.49 (eBook)
Niles Yan Zhao
This book review examines Phakiti's 2025 work on language testing and assessment, which bridges theoretical foundations with practical applications for TESOL professionals.
- PaperRELC Journal7 May 2026
Assessing the impact of the Common European Framework of Reference for Languages on policy and testing practices in Thai higher education
Anchana Rukthong, Punjaporn Pojanapunya, Somruedee Khongput
This study analyzed the implementation of CEFR-based English Exit Exams at 12 Thai universities, finding inconsistent skill coverage and unclear alignment of scores with CEFR levels, which led to doubts about exam quality among practitioners. While 84% of surveyed stakeholders supported the policy for raising learner awareness and motivation, opponents questioned whether test scores accurately reflected actual communicative ability. The authors conclude that English Exit Exams alone may not effectively drive CEFR-informed classroom practice.
Original abstract
The application and impact of the Common European Framework of Reference for Languages have been observed worldwide, especially in testing practices, since its introduction in 2001. This is partly because Common European Framework of Reference-based policies have often used the framework's proficiency scales as a key indicator of success. Although the Common European Framework of Reference as policy is expected to improve current practices, there appear potential challenges that could undermine its effectiveness in particular contexts. This study reports on the case of the Common European Framework of Reference-based policy in Thai higher education, where an English Exit Exam has been imposed to verify the proficiency of graduates. To obtain a clear picture of how the policy has been implemented, the study analysed 12 sets of English Exit Exam materials from 12 institutions, 44 questionnaire responses from university lecturers and administrators, and 17 online interviews with practitioners who were involved in the development of English Exit Exams. The results show that the English Exit Exams of different institutions focused on assessing different skills, sub-skills and linguistic knowledge, with exams including listening, reading, vocabulary and grammar sections; however, how the exam results were aligned with the Common European Framework of Reference scales remained unclear despite the framework's centrality in policy, and this made about half of the interview participants doubt the quality of the English Exit Exam. Eighty-four percent of the questionnaire respondents supported the policy, considering it a tool for raising learners’ awareness and motivation in learning English. The opponents, however, question the effectiveness of the English Exit Exam practices, being concerned that the test scores may not accurately indicate actual ability in English communication. The study concludes that although a test can theoretically serve as a mechanism to impose language policy, the endorsement of English Exit Exams alone may not effectively function as an instrument to drive the Common European Framework of Reference-informed practice at a classroom level.
- PaperLanguage Testing26 Apr 2026
Justifying the Score or Informing the Stakeholder? Transparency Challenges in Large-Scale Language Testing
Vahid Aryadoust
Large-scale language testing needs greater transparency, according to this viewpoint, which contrasts transparency with persuasiveness in argument-based validity. Two measures are proposed: publishing test-form-specific validity reports and clearly explaining test limitations, inspired by pharmaceutical industry norms. The goal is to shift emphasis from justification to transparency and align with open science.
Original abstract
Language testing is both an evaluative practice and a commercial enterprise shaped by market forces. Within this context, test developers have a responsibility to ensure transparency with test users, particularly when scores inform high-stakes decisions. This Viewpoint contrasts transparency in communicating the truth about language tests with the persuasiveness of argument-based validity, noting that persuasiveness, although central to such arguments, is not equivalent to transparency. Two measures are proposed to strengthen transparency. First, test developers should publish test-form-specific validity reports detailing content, psychometric properties, and interpretation guidelines for each form. Second, they should clearly explain a test’s limitations to the public, especially when scores are used in high-stakes settings, such as immigration or university admission, without adequate validation. The latter measure draws on regulatory norms in the pharmaceutical industry, where transparency can protect consumers from potential misuse. Specific steps are outlined to support these measures. Overall, these proposals aim to shift the emphasis from justification and persuasion toward transparency and align language testing practices more closely with open science principles.
- PaperETS Research Report Series1 Apr 2026
Beyond Score Correlations: A Content Comparison of IELTS Academic and TOEFL iBT® Tests in the Context of a Score Concordance Study
Sara Cushing
This study compares the content of IELTS Academic and TOEFL iBT across reading, listening, writing, and speaking, finding substantial construct overlap in reading and writing but notable differences in listening and speaking, such as TOEFL's emphasis on academic content and integrated skills versus IELTS's general listening and examiner-mediated speaking. Overall, the results support score concordance tables while cautioning against treating scores as fully interchangeable.
Original abstract
Score concordance tables are widely used by higher education institutions to compare scores from different English language proficiency tests, yet their validity depends on the extent to which the tests measure comparable constructs. This study examines the content comparability of the International English Language Testing System (IELTS) Academic and TOEFL iBT® tests in the context of a recent co-sponsored score concordance study. Moving beyond score correlations, the analysis compares test content across the four language skills—reading, listening, writing, and speaking—using published research, official test documentation, and publicly available sample materials. The comparison is framed by the validity frameworks adopted by each test provider and focuses on task characteristics, linguistic demands, response formats, and scoring criteria. Results indicate substantial overlap in the constructs assessed by both tests, particularly in reading and writing, where tasks target similar academic language skills and are evaluated using comparable criteria. More pronounced differences emerge in listening and speaking, with TOEFL iBT placing greater emphasis on academic content, integrated skills, and pragmatic inference, while IELTS includes more general listening contexts and examiner-mediated interaction in speaking. Despite these differences, both tests provide multiple opportunities for test takers to engage with extended discourse and demonstrate receptive and productive language abilities. Overall, the findings support the use of score concordance tables between IELTS Academic and TOEFL iBT, while emphasizing the need for cautious interpretation and recognition that scores should not be treated as fully interchangeable. Suggested citation: Cushing, S. T. (2026). Beyond score correlations: A content comparison of IELTS Academic and TOEFL iBT® tests in the context of a score concordance study (TOEFL Research Report No. RR-107). ETS. https://doi.org/10.64634/qercg225
- PaperLanguage Testing27 Mar 2026
The 2001 ILTA Code of Ethics: Philosophical foundations and historical context
Bart Deygers, Daan Van Cauwenberge
This paper examines the philosophical foundations and historical context of the ILTA Code of Ethics, analyzing its strategic rationale and unresolved tensions. It aims to deepen understanding of ethical frameworks in language testing as the association revises the code.
Original abstract
While ethical and moral philosophy underpins core concepts in language testing, few publications in the field engage with moral philosophy in a sustained or systematic way. A notable exception is the International Language Testing Association’s (ILTA) Code of Ethics (COE), ratified in 2000, which continues to guide professional practice. As ILTA initiates a revision of the COE, this paper examines the objectives and philosophical foundations of the original COE. Drawing on the document itself and key writings by its principal author, Alan Davies, and contemporary authors, we offer an analysis that reveals both the strategic rationale behind the Code and the unresolved tensions within it. Our aim is to deepen the field’s understanding of the ethical frameworks that have shaped the language testing profession as ILTA traces a way forward in revising the COE.
- PaperLanguage Testing27 Mar 2026
How test-taker level of schooling interacts with pass probability on high-stakes tests for citizenship
Marieke Vanbuel, Edit Bugge
This study analyzes test results from 79,794 migrants in Norway, revealing that learners with limited formal education in their L1 have significantly lower pass probabilities on citizenship tests, especially the knowledge of society (KoS) test, which functions as an implicit language test. The impact of L2 proficiency on passing KoS tests is stronger for less educated learners, highlighting barriers for adult L2 learners with limited schooling.
Original abstract
In most Western European countries, aspiring citizens must pass both a second language (L2) test and a knowledge of society (KoS) test, often administered in the L2. While several studies highlighted the social inequalities such tests may create, few have quantified these disparities. This study, anchored in the IMPECT project, analyses the test results of 79,794 migrants in Norway to explore the implications of these requirements, particularly for adult L2 learners with limited formal education in their first language (L1). Logistic regression analyses reveal that (1) learners with limited formal L1 education have significantly lower probabilities of passing the tests compared to those with higher education; (2) disparities are greater for KoS tests than for the oral A2 language test; (3) L2 scores strongly predict the likelihood of passing KoS tests, substantiating the assumption that KoS tests are implicit language tests; and (4) the impact of language skills on passing KoS tests varies by schooling level, with L2 scores being a stronger predictor for learners with limited education than for those with tertiary degrees. These findings highlight barriers faced by learners with limited L1 schooling in the citizenship acquisition process and underscore the need for targeted policy and support mechanisms.
- PaperETS Research Report Series9 Oct 2025
TOEFL iBT® Technical Manual
Venessa Manna, Shuhong Li, Spiros Papageorgiou, Lixiong Gu
This technical manual details the design, scoring, and intended uses of the TOEFL iBT test, targeting test-takers and language use domains. It presents a research agenda to support score interpretation and validity evidence, and is designed as a living document to be updated as the test evolves, especially with changes starting January 2026.
Original abstract
This technical manual describes the purpose and intended uses of the TOEFL iBT test, its target test-taker population, and relevant language use domains. The test design and scoring procedures are presented first, followed by a research agenda intended to support the interpretation and use of test scores. Given the updates to the test starting January 2026, this technical manual is intended to serve as an overview and rationale for the test design as well as a reference point for informing investigations of validity evidence to support the intended test uses over time. Designed as a living document, this manual will be updated as the test's design, administration, scoring, and evidence of measurement quality (including reliability, validity, and fairness) evolve, along with its intended uses. Suggested citation: Manna, V. F., Li, S., Papageorgiou, S., & Gu, L. (2025). TOEFL iBT® technical manual (TOEFL Research Report No. RR-106). ETS. https://doi.org/10.64634/eje8f497