item-response-theory
Filtering by topic item-response-theory(2)Clear all filters
- PaperComputers and Education: Artificial Intelligence17 Jun 2026
The impact of item-writing flaws on difficulty and discrimination in item response theory
Robin Schmucker, Steven Moore
This study analyzed 7,126 multiple-choice questions across STEM subjects using an automated item-writing flaw rubric. Significant links were found between the number of flaws and IRT difficulty/discrimination parameters, especially in life/earth and physical sciences. The findings support using automated IWF analysis for initial item screening.
Original abstract
High-quality test items are essential for educational assessments, particularly within Item Response Theory (IRT). Traditional validation methods rely on resource-intensive pilot testing to estimate item difficulty and discrimination. More recently, Item-Writing Flaw (IWF) rubrics emerged as a domain-general approach for evaluating test items based on textual features. This method offers a scalable, pre-deployment evaluation without requiring student data, but its predictive validity concerning empirical IRT parameters is underexplored. To address this gap, we conducted a study involving 7,126 multiple-choice questions across various STEM subjects (physical science, mathematics, and life/earth sciences). Using an automated approach, we annotated each question with a 19-criteria IWF rubric and studied relationships to data-driven IRT parameters. Our analysis revealed statistically significant links between the number of IWFs and IRT difficulty and discrimination parameters, particularly in life/earth and physical science domains. We further observed how specific IWF criteria can impact item quality more and less severely (e.g., negative wording vs. implausible distractors) and how they might make a question more or less challenging. Overall, our findings establish automated IWF analysis as a valuable supplement to traditional validation, providing an efficient method for initial item screening, particularly for flagging low-difficulty MCQs. Our findings show the need for further research on domain-general evaluation rubrics and algorithms that understand domain-specific content for robust item validation.
- PaperETS Research Report Series19 Nov 2025
An Evaluation of Item Fit Based on Generalized Residual Item Response Functions
Xiangyi Liao, Peter Van Rijn, Sandip Sinharay
This study develops a new method for evaluating item fit in item response theory (IRT) models by summarizing generalized residuals into a single statistic per item, accounting for estimation error. Simulations show similar Type I error rates to an existing method with slight improvements for small samples, but low power except for problematic items.
Original abstract
Evaluation of item ft for item response theory (IRT) models often involves a comparison of the observed and expected item response functions (IRFs). Several statistics have been suggested for evaluating item ft based on the discrepancy between IRFs, but the asymptotic distributions of the statistics under the null hypothesis are often not well established. Haberman et al. developed a method for evaluating the ft of IRFs based on generalized residuals. These residuals are functions of the latent proficiency variable in the IRT model and follow the standard normal distribution asymptotically. We develop a method to summarize these generalized residuals into a single summary statistic for each item and evaluate its asymptotic distribution. Kondratek suggested a similar Wald-type statistic, but without accounting for the uncertainty in the estimation of the item parameters. Our method combines the work of Haberman and Kondratek, resulting in a single ft statistic per item while accounting for estimation error. A series of simulations was carried out to investigate the performance of our statistic and compare it to several popular item ft statistics. Our method resulted in similar Type I errors as Kondratek’s statistic, with slightly better results in the case of small samples. Furthermore, the recovery was consistent across different levels of item difficulty, and power of the new item ft statistic was relatively low, except for problematic individual items, but this result was found with two competing statistics as well. Suggested citation: Liao, X., van Rijn, P., & Sinharay, S. (2025). An evaluation of item fit based on generalized residual item response functions (Research Report No. RR-25-13). ETS. https://doi.org/10.64634/b68vz316