ai-assessment
Filtering by topic ai-assessment(3)Clear all filters
- PaperEdArXiv (OSF Preprints)20 Jul 2026
Learning Orchestration: What Can Count as Learning in AI-Mediated Assessment?
James Wood
Proposes Learning Orchestration as a framework for credible AI-mediated assessment, arguing that evidence of learning must consider how learners coordinate cognitive, social, and technological resources, not just final artifacts. Distinguishes independent capability, AI-supported orchestration, and accountable verification as distinct claims requiring different evidence.
Original abstract
Responses to generative AI in higher education assessment have focused on the architecture of AI use: secured and open task designs, permission scales, declarations and process evidence. These approaches specify where and how AI may be used, but leave two harder questions unresolved: what capability an assessment claims the learner has, and what evidence could warrant that claim. This article argues that credible evidence of learning in AI-mediated assessment cannot be inferred from the artefact alone, but depends partly on the quality of the learner’s orchestration of the resources that shape it. Learning Orchestration is defined as the purposeful and accountable coordination of distributed cognitive, social, informational and technological resources, through cognitive-metacognitive, epistemic and ethical judgement, to advance learning, develop disciplinary judgement and produce work, decisions or claims that the learner can explain, justify and defend. The article distinguishes Learning Orchestration from self-regulated learning, evaluative judgement, AI literacy, epistemic agency, academic integrity and assessment validity. It contributes a learner-side account linking the quality of resource coordination to the claims about learning that assessment evidence can warrant. The resulting Assessment Claims Framework distinguishes independent capability, AI-supported orchestration and accountable verification of AI-shaped work as claims requiring different forms of evidence. Without these distinctions, universities risk certifying fluent AI-shaped performance while overstating what learners understand, can do independently or can responsibly verify and defend. Credible AI-mediated assessment therefore depends not only on regulating AI use, but on teaching, eliciting and judging the quality of learners’ orchestration.
- PaperJournal of Second Language Writing14 May 2026
Evaluating ChatGPT-4o as an AI assessor in Chinese as a Second Language writing: Reliability through generalizability theory, feedback actionability, and teacher-student perceptions
Xiaosheng Zhou, Hanwei Wu, Ying Soon Goh
This study evaluates ChatGPT-4o's reliability as an AI assessor for Chinese as a Second Language writing using generalizability theory, and examines feedback actionability and teacher-student perceptions.
- PaperERIC — ELT & TESOL1 Jan 2025
Assessing the Accuracy of Automated Writing Evaluation in Predicting English Language Arts Proficiency for Middle-Grade English Language Learners and Non-English Language Learners
Fan Zhang, Joshua Wilson
This study examined the accuracy of MI Write, an automated writing evaluation system, in predicting non-proficiency on the Smarter Balanced ELA assessment for middle-grade students, considering grade level, seasonality, and language status. Results showed acceptable overall classification accuracy, strongest for non-ELLs and Grade 7 students, but more inconsistent for English language learners. The findings highlight the potential and limitations of AI-based assessment systems for diverse student populations.
Original abstract
This study examines the accuracy of fall, winter, and spring benchmark writing assessments, scored by the MI Write automated writing evaluation system, for predicting non-proficiency on the Smarter Balanced ELA assessment. This study considers how grade level, seasonality, and language status influence classification accuracy using Receiver Operating Characteristic (ROC) curve analyses. The results indicate that MI Write demonstrated acceptable overall classification accuracy, with the strongest performance among non-ELLs and Grade 7 students. However, accuracy was more inconsistent for ELLs. Across all grades and subgroups, the d-based cutpoints consistently provided the best balance between sensitivity and specificity. Implications for adopting AI-based assessment systems within the middle grades are discussed.