scoring-consistency
Filtering by topic scoring-consistency(2)Clear all filters
- PaperLanguage Teaching Research16 Jul 2026
Comparing Teacher and Artificial Intelligence Scoring in Writing Assessment: A Generalizability Theory Analysis
Burak Asma
Compared teacher and AI scoring of middle school essays with and without rubrics using generalizability theory. AI tools showed higher consistency and better differentiation of individual differences, while teachers were influenced by biases and mood. Teachers acknowledged the potential of rubrics and AI feedback for more consistent results.
Original abstract
This study examined the use of artificial intelligence tools, which have garnered significant attention in recent years, in the assessment and evaluation processes of language education. For this purpose, student essays were scored by Turkish middle school teachers and artificial intelligence tools both with and without the use of a rubric, and the findings were evaluated based on generalizability theory. Additionally, the research findings were shared with participants to gather qualitative data, which were analysed using the inductive thematic analysis method to support the research results. The findings revealed that in evaluations conducted without a rubric, teachers were limited in their ability to distinguish individual differences and demonstrated low scoring consistency. In contrast, artificial intelligence tools were more effective in distinguishing individual differences and exhibited high consistency. In evaluations conducted using a rubric, scoring consistency increased in both groups, although, as in the first evaluation, artificial intelligence tools demonstrated a higher level of consistency. Regarding the research findings, teachers expressed that individual biases, mood, and professional experiences influenced their scoring processes and emphasized the potential of rubrics and artificial intelligence-supported feedback systems for achieving more consistent results. Artificial intelligence tools, on the other hand, highlighted their independence from subjective factors but stressed the need for more diverse and generalizable datasets to further enhance their evaluation capacities.
- PaperETS Research Report Series31 Dec 2025
Using Ordinal Rescore Measures to Monitor Rater Drift
John Donoghue, Adrienne Sgammato
Ordinal rescore measures are used to monitor rater drift in constructed response scoring across occasions. An alternative analysis that explicitly conditions on the rescore design is contrasted with usual trend analysis and found to be effective. Omnibus measures based on t-tests showed marginally higher power than those based on d-statistics in detecting drift.
Original abstract
When constructed response items are used on more than one occasion, a natural concern is whether the scoring is consistent (e.g., not more lenient or strict) across the occasions. It is common to conduct trend scoring, in which a set of Occasion A responses are rescored at Occasion B. The responses are usually selected according to some rescore design, such as being balanced (with an equal number from each score category), proportional to the distribution of Occasion A scores, or a mixed version of these two designs. Recent work has demonstrated that treating the two-way table as if it arose from multinomial sampling is incorrect and can yield seriously biased estimates of whether the scores are lower or higher at Occasion B. The present study builds on these results by incorporating ordinal measures of change. It contrasts the usual trend analysis with an alternative analysis that explicitly conditions on the rescore design and finds only the latter to be effective. Omnibus measures based on combining the individual t-tests or d-statistics are examined. Measures were somewhat conservative in Type I error control and had good power to detect drift. Omnibus measures based on t-tests had marginally higher power, having higher correct detection rates than those based on the d-statistic in 1%–8% of the cases. The difference between the best versions (E weighted, which is based on t-tests, vs. D weighted, which is based on d-statistics) was only 1.8%. Suggested citation: Donoghue, J. R., & Sgammato, A. (2025). Using ordinal rescore measures to monitor rater drift(Research Report No. RR-25-15). ETS.