automated-writing-evaluation
Filtering by topic automated-writing-evaluation(4)Clear all filters
- PaperAssessing Writing23 Jun 2026
From keystrokes to scores: Toward a multidimensional predictive model of writing evaluation by humans and large language models across linguistic, cognitive, and social dimensions
Qiao Gan, Benjamin Adams
Analyzed 4245 argumentative essays with keystroke and demographic data to compare how linguistic, cognitive, and social factors covary with essay scores from human raters and large language models (LLMs). Found moderate agreement between human and LLM scores but different patterns of association with these factors, indicating that human and machine assessments rely on partially different cues.
Original abstract
Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.
- PaperLanguage Testing25 Mar 2026
™ChatGPT for automated writing evaluation: Scoring and feedback across prompt conditions
Yewon Lee, Myunghwan Hwang
Six ChatGPT models with different prompt configurations were tested against human raters on 60 EFL writing samples. Prompt design significantly influenced scoring consistency and severity, with Chain-of-Thought and Fill-in-the-blank prompts yielding higher reliability. Learners perceived the feedback positively, but reasoning-intensive domains still required human oversight.
Original abstract
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
- PaperReCALL29 Dec 2025
Impact of prompt sophistication on ChatGPT’s output for automated written corrective feedback
Na Luo, Yifan Wang, Zhe (Victor) Zhang, Yile Zhou et al.
Compared ChatGPT's automated written corrective feedback (AWCF) accuracy across generic and domain-specific prompts against Grammarly. Found that domain-specific prompts, especially one-shot, significantly improved error detection, with zero-shot matching Grammarly and one-shot surpassing it. However, even the most sophisticated prompt still showed limitations compared to Grammarly.
Original abstract
The emergence of large language models, exemplified by ChatGPT, has garnered growing attention for their potential to generate feedback in second language writing, particularly automated written corrective feedback (AWCF). In this study, we examined how prompt design – a generic prompt and two domain-specific prompts (zero-shot and one-shot) enriched with comprehensive domain knowledge about written corrective feedback (WCF) – influences ChatGPT’s ability to provide AWCF. The accuracy and coverage of ChatGPT’s feedback across these three prompts were benchmarked against Grammarly, a widely used traditional automated writing evaluation (AWE) tool. We find that ChatGPT’s ability in flagging language errors grew considerably with prompt sophistication driven by the integration of domain-specific knowledge and examples. While the generic prompt resulted in substantially lower performance than Grammarly, the zero-shot prompt achieved comparable results to it and the one-shot prompt surpassed it considerably in error detection. Notably, the most pronounced improvement in ChatGPT’s performance was observed in its detection of frequent error categories, including those of word choice or expression, direct translation, sentence structure and pronoun. Nonetheless, even with the most sophisticated prompt, ChatGPT still displayed certain limitations when compared to Grammarly. Our study has both theoretical and practical implications. Theoretically, it lends empirical evidence to Knoth et al .’s (2024) proposition to separate domain-specific AI literacy from generic AI literacy. Practically, it sheds light on the pedagogical application and technical development of AWE systems.
- PaperERIC — ELT & TESOL1 Jan 2025
Assessing the Accuracy of Automated Writing Evaluation in Predicting English Language Arts Proficiency for Middle-Grade English Language Learners and Non-English Language Learners
Fan Zhang, Joshua Wilson
Examines the accuracy of MI Write, an automated writing evaluation system, in predicting non-proficiency on the Smarter Balanced ELA assessment for middle-grade students. Results show acceptable overall accuracy, stronger for non-ELLs and Grade 7, but less consistent for ELLs. d-based cutpoints offered best sensitivity-specificity balance across subgroups.
Original abstract
This study examines the accuracy of fall, winter, and spring benchmark writing assessments, scored by the MI Write automated writing evaluation system, for predicting non-proficiency on the Smarter Balanced ELA assessment. This study considers how grade level, seasonality, and language status influence classification accuracy using Receiver Operating Characteristic (ROC) curve analyses. The results indicate that MI Write demonstrated acceptable overall classification accuracy, with the strongest performance among non-ELLs and Grade 7 students. However, accuracy was more inconsistent for ELLs. Across all grades and subgroups, the d-based cutpoints consistently provided the best balance between sensitivity and specificity. Implications for adopting AI-based assessment systems within the middle grades are discussed.