llm-assessment
Filtering by topic llm-assessment(2)Clear all filters
- PaperAssessing Writing18 Jul 2026
Modeling the reading-to-writing pipeline: Knowledge graph and LLM-based assessment framework for source-based writing
Byungyeon Yun, Miranda Moe, Lauren E. Flynn, Püren Öncel et al.
A framework using knowledge graphs and large language models assesses source-based writing by modeling the reading-to-writing pipeline.
- PaperAssessing Writing23 Jun 2026
From keystrokes to scores: Toward a multidimensional predictive model of writing evaluation by humans and large language models across linguistic, cognitive, and social dimensions
Qiao Gan, Benjamin Adams
This study examines how linguistic, cognitive, and social factors covary with essay scores from human raters and large language models (LLMs), using keystroke-logging data and demographic metadata. Moderate agreement was found between human and LLM scores, but the two scoring systems showed different patterns of association with these factors, indicating reliance on partially different cues.
Original abstract
Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.