educational-assessment
Filtering by topic educational-assessment(4)Clear all filters
- PaperarXiv — AI in Education (cs.CY)15 Jul 2026
When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring
Nischal Ashok Kumar, Payu Wittawatolarn, Sana Kang, Marisa C. Peczuh et al.
The paper investigates cross-rubric generalization in automated essay scoring, where models trained on essays scored under one rubric must perform on new rubrics targeting different aspects. Using a trait-based intermediate representation and target-essay supervision, the approach improves macro F1 by 5% in the hardest setting. Their best open-source Llama-based model outperforms GPT-5-mini prompting by 2.1% and trails GPT-5 by 1.9%.
Original abstract
Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target essays are unseen during training. We further find that increasing target-essay supervision improves performance, with our best fine-tuned open-source Llama-based model outperforming GPT-5-mini prompting by 2.1% macro F1 and trailing GPT-5 by 1.9%. These results show that trait-based intermediate structure and controlled supervision improve generalization to unseen rubrics.
- PaperAssessment & Evaluation in Higher Education1 Jul 2026
Which grades predict what? A more nuanced understanding of using high school results for university admission
Sebastiaan Steenman, Ada Kool
This study examined the predictive value of high school grades for university performance across different cognitive learning objectives, programs, and time. Overall high school GPA consistently outperformed subject-specific grades, but taking related high school subjects was associated with small improvements. Predictive strength declined over the course of bachelor programs and was better for lower-order cognitive skills.
Original abstract
While high school grades are widely used for university admissions, little is known about which specific high school grades best predict what type of performance at university. This study examines the predictive value of overall high school GPA (grade point average), grades for subsets of subjects, and the added value of having taken specific subjects, for university performance across different cognitive learning objectives, different programmes and over time. Using data from multiple cohorts of six undergraduate programmes at a large Dutch research university, we show that the overall high school GPA consistently outperforms subsets of discipline-related subjects, suggesting that high school grades primarily represent general learning skills and traits. However, having taken a specific related high school subject was generally associated with better university performance, although effect sizes were small. High school grades predicted performance better on assessments targeting lower-order cognitive skills than complex academic tasks. No significant differences emerged between the predictive value of high school final-year and penultimate-year grades. Finally, the predictive strength declined over the course of the three-year bachelor programmes. These findings highlight the need for careful consideration of which high school grades to use in admissions and provide practical suggestions for university admissions officers to do so.
- PaperComputers and Education: Artificial Intelligence17 Jun 2026
The impact of item-writing flaws on difficulty and discrimination in item response theory
Robin Schmucker, Steven Moore
This study analyzed 7,126 multiple-choice questions across STEM subjects using an automated item-writing flaw rubric. Significant links were found between the number of flaws and IRT difficulty/discrimination parameters, especially in life/earth and physical sciences. The findings support using automated IWF analysis for initial item screening.
Original abstract
High-quality test items are essential for educational assessments, particularly within Item Response Theory (IRT). Traditional validation methods rely on resource-intensive pilot testing to estimate item difficulty and discrimination. More recently, Item-Writing Flaw (IWF) rubrics emerged as a domain-general approach for evaluating test items based on textual features. This method offers a scalable, pre-deployment evaluation without requiring student data, but its predictive validity concerning empirical IRT parameters is underexplored. To address this gap, we conducted a study involving 7,126 multiple-choice questions across various STEM subjects (physical science, mathematics, and life/earth sciences). Using an automated approach, we annotated each question with a 19-criteria IWF rubric and studied relationships to data-driven IRT parameters. Our analysis revealed statistically significant links between the number of IWFs and IRT difficulty and discrimination parameters, particularly in life/earth and physical science domains. We further observed how specific IWF criteria can impact item quality more and less severely (e.g., negative wording vs. implausible distractors) and how they might make a question more or less challenging. Overall, our findings establish automated IWF analysis as a valuable supplement to traditional validation, providing an efficient method for initial item screening, particularly for flagging low-difficulty MCQs. Our findings show the need for further research on domain-general evaluation rubrics and algorithms that understand domain-specific content for robust item validation.
- PaperLanguage Testing10 Jun 2026
Beyond Traditional Differential Item Functioning Detection: A Rasch Tree Approach to Evaluating Item Fairness in a Large-Scale German Reading Comprehension Test
Farshad Effatpanah, Olga Kunina-Habenicht, Katharina Antonia Michiko Tremmel, Philipp Sonnleitner
Applied the Rasch tree model to detect differential item functioning (DIF) in a large-scale German reading comprehension test for fifth graders, using covariates such as gender, socioeconomic status, immigration status, and personality. The analysis identified eleven items with moderate to large DIF across four splitting nodes, showing that combinations of covariates affected test performance without the need for pre-specified groups.
Original abstract
A key psychometric phenomenon in educational testing is differential item functioning (DIF), which evaluates whether specific test items function differently across subgroups of examinees who have the same level of the underlying (latent) ability. DIF happens when examinees with the same latent trait have different probabilities of correctly responding to a test item, influenced by their subgroup membership. This study aims to illustrate the application of the recursive partitioning Rasch tree model, shortly called the Rasch tree, to examine DIF in a large-scale German reading comprehension test for fifth-grade elementary students across gender, socioeconomic status, immigration status, personality disposition, and the need for cognition. Unlike the conventional DIF detection methods, the Rasch tree does not require a pre-specification of groups for exploring DIF, and continuous covariates can be easily included in the analysis. To investigate DIF of the test, item responses of 4,252 students were analyzed. The Rasch tree analysis generated nine non-predefined nodes, with slightly different patterns of item difficulties. Eleven items were flagged as exhibiting moderate and large DIF in the four splitting nodes. No splits were produced based on the need for cognition. The results indicated that the combination of the covariates impacted students’ test performance.