test-fairness
Filtering by topic test-fairness(3)Clear all filters
- PaperTESOL Quarterly10 Jul 2026
Handwritten Versus Typed Notes: The Impact of Note‐Taking Modes in Second Language Listening Tests
Jieun Kim
A study with 305 L2-English adults found no significant overall score differences between handwritten, typed, and no note-taking conditions on TOEFL iBT lecture comprehension tasks. However, item-level differential functioning and note content differences emerged, suggesting test developers should reconsider note-taking policies for validity and fairness. Pedagogical implications include allowing students autonomy in choosing note-taking modes and providing instruction in various strategies.
Original abstract
Technological advancements have influenced note‐taking practices in classrooms as well as their treatment in standardized language assessments. Many high‐stakes English listening tests (e.g., TOEFL iBT, IELTS) include varying note‐taking guidelines, often established without empirical support. This study examined the effects of note‐taking modes on listening performance and note content. A total of 305 L1‐Korean L2‐English adults were randomly assigned to handwriting ( n = 102), typing ( n = 102), or no note‐taking ( n = 101) conditions and completed TOEFL iBT lecture comprehension tasks. Linear regression results revealed no significant differences in overall test scores across modes. Uniform Differential Item Functioning (DIF) analyses indicated that one item favored handwriting and another favored typing, while non‐uniform DIF emerged for the same items only among lower‐ability participants. When examining note‐taking features, no significant differences were found in word count or translanguaging across modes. However, handwritten notes contained more information units, verbatim transcription, and nonlinguistic elements. The overall lack of significant performance differences, together with finer‐grained differences at the item and note‐content levels, calls for test developers to reconsider note‐taking policies in terms of test validity and fairness. Pedagogical implications are discussed, highlighting the importance of granting students autonomy to choose note‐taking modes and providing instructional sessions to experience various note‐taking strategies.
- PaperLanguage Testing10 Jun 2026
Beyond Traditional Differential Item Functioning Detection: A Rasch Tree Approach to Evaluating Item Fairness in a Large-Scale German Reading Comprehension Test
Farshad Effatpanah, Olga Kunina-Habenicht, Katharina Antonia Michiko Tremmel, Philipp Sonnleitner
Applied the Rasch tree model to detect differential item functioning in a large-scale German reading comprehension test for fifth graders, analyzing responses from 4,252 students across gender, socioeconomic status, immigration status, and personality. The method identified four splitting nodes with eleven items showing moderate to large DIF without requiring pre-specified groups, revealing how covariate combinations impact test performance.
Original abstract
A key psychometric phenomenon in educational testing is differential item functioning (DIF), which evaluates whether specific test items function differently across subgroups of examinees who have the same level of the underlying (latent) ability. DIF happens when examinees with the same latent trait have different probabilities of correctly responding to a test item, influenced by their subgroup membership. This study aims to illustrate the application of the recursive partitioning Rasch tree model, shortly called the Rasch tree, to examine DIF in a large-scale German reading comprehension test for fifth-grade elementary students across gender, socioeconomic status, immigration status, personality disposition, and the need for cognition. Unlike the conventional DIF detection methods, the Rasch tree does not require a pre-specification of groups for exploring DIF, and continuous covariates can be easily included in the analysis. To investigate DIF of the test, item responses of 4,252 students were analyzed. The Rasch tree analysis generated nine non-predefined nodes, with slightly different patterns of item difficulties. Eleven items were flagged as exhibiting moderate and large DIF in the four splitting nodes. No splits were produced based on the need for cognition. The results indicated that the combination of the covariates impacted students’ test performance.
- PaperERIC — Assessment & second language1 Jan 2025
A Systematic Review of Differential Item Functioning in Second Language Assessment
Xueliang Chen, Vahid Aryadoust, Wenxin Zhang
This systematic review of 83 articles on differential item functioning (DIF) in second language assessments found that classical methods like Rasch and Mantel-Haenszel are dominant, while emerging cognitive diagnostic models are also used. Most studies examine gender and language background as grouping variables, focus on receptive skills, and rely on speculative rather than empirical justifications for DIF causes. The review calls for improved DIF practices, alternative methods aligning with modern views of bias, and better accounting for test taker diversity to enhance fairness and validity.
Original abstract
The growing diversity among test takers in second or foreign language (L2) assessments makes the importance of fairness front and center. This systematic review aimed to examine how fairness in L2 assessments was evaluated through differential item functioning (DIF) analysis. A total of 83 articles from 27 journals were included in a systematic review. The findings suggested that classical DIF techniques were dominant in use, particularly Rasch-based methods, the Mantel-Haenszel procedure, item response theory (IRT) approaches, logistic regression, and SIBTEST, but emerging methods such as DIF analysis based on cognitive diagnostic models were also identified. Most DIF studies examined manifest grouping variables such as gender and language background and were based on assessments of receptive language skills such as reading and listening comprehension. DIF analyses were mostly conducted in an exploratory fashion and causes of DIF were often justified on speculative rather than empirical grounds. In addition, the quality of DIF analyses was undermined by suboptimal reporting practices. Our results suggest the need to improve current DIF practices, to consider alternative DIF detection methods aligning with emerging views of measurement bias, and to adequately account for the heterogeneity of L2 test takers. The findings have implications for test design and use, fairness, and validity in L2 assessments.