automated-writing-evaluation
Filtering by topic automated-writing-evaluation(3)Clear all filters
- PaperAssessing Writing23 Jun 2026
From keystrokes to scores: Toward a multidimensional predictive model of writing evaluation by humans and large language models across linguistic, cognitive, and social dimensions
Qiao Gan, Benjamin Adams
Analyzed 4245 essays with keystroke and demographic data to compare human and LLM scoring across linguistic, cognitive, and social dimensions. Found moderate agreement but different patterns of association, indicating that human and LLM assessments rely on partially distinct cues.
Original abstract
Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.
- PaperLanguage Testing25 Mar 2026
™ChatGPT for automated writing evaluation: Scoring and feedback across prompt conditions
Yewon Lee, Myunghwan Hwang
ChatGPT's performance as an Automated Writing Evaluation (AWE) system was tested by comparing its scoring with human raters on 60 EFL essays from Korean university students. Prompt design significantly affected scoring consistency, with Chain-of-Thought reasoning combined with Fill-in-the-blank scaffolding yielding higher reliability, though domain-level bias remained uncontrolled. Learners generally perceived ChatGPT's feedback positively, suggesting prompt-calibrated AWE can serve as a supplementary tool for writing assessment.
Original abstract
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
- PaperReCALL29 Dec 2025
Impact of prompt sophistication on ChatGPT’s output for automated written corrective feedback
Na Luo, Yifan Wang, Zhe (Victor) Zhang, Yile Zhou et al.
The study examined how prompt design (generic vs. domain-specific zero-shot and one-shot) affects ChatGPT's accuracy and coverage in providing automated written corrective feedback (AWCF), benchmarked against Grammarly. Domain-specific prompts, especially one-shot, significantly improved error detection, surpassing Grammarly in some categories. However, ChatGPT still showed limitations compared to Grammarly.
Original abstract
The emergence of large language models, exemplified by ChatGPT, has garnered growing attention for their potential to generate feedback in second language writing, particularly automated written corrective feedback (AWCF). In this study, we examined how prompt design – a generic prompt and two domain-specific prompts (zero-shot and one-shot) enriched with comprehensive domain knowledge about written corrective feedback (WCF) – influences ChatGPT’s ability to provide AWCF. The accuracy and coverage of ChatGPT’s feedback across these three prompts were benchmarked against Grammarly, a widely used traditional automated writing evaluation (AWE) tool. We find that ChatGPT’s ability in flagging language errors grew considerably with prompt sophistication driven by the integration of domain-specific knowledge and examples. While the generic prompt resulted in substantially lower performance than Grammarly, the zero-shot prompt achieved comparable results to it and the one-shot prompt surpassed it considerably in error detection. Notably, the most pronounced improvement in ChatGPT’s performance was observed in its detection of frequent error categories, including those of word choice or expression, direct translation, sentence structure and pronoun. Nonetheless, even with the most sophisticated prompt, ChatGPT still displayed certain limitations when compared to Grammarly. Our study has both theoretical and practical implications. Theoretically, it lends empirical evidence to Knoth et al .’s (2024) proposition to separate domain-specific AI literacy from generic AI literacy. Practically, it sheds light on the pedagogical application and technical development of AWE systems.