reliability-validity
Filtering by topic reliability-validity(1)Clear all filters
- PaperAssessing Writing26 Jun 2026
Assessing the reliability and validity of large language models in automatic essay scoring
SCOTT CROSSLEY, Langdon Holmes, Wesley Morris
Large language model (LLM) based automated essay scoring (AES) systems were developed and tested for reliability, agreement with human raters, and convergent validity using linguistic components. Both representation and generative LLM AES systems showed strong reliability and agreement, but the representation system exhibited differential correlations with human scores regarding text length and type-token ratio, while the generative system showed no differences. The findings support the use of LLM-based AES in standardized writing assessments for secondary students, though further research is needed.
Original abstract
With the advent of artificial intelligence, large language model (LLM) based Automated Essay Scoring (AES) systems have been developed that can consistently make human-like decisions that do not depend fully on surface level linguistic features. However, research into the use of LLM-based AES systems is limited and little is known about the reliability, agreement, or validity of the systems. The goal of this study was to provide evidence for the reliability, agreement, and validity of LLM-based AES systems in a standardized writing assessment used for secondary school students. Both representation and generative LLM-based AES systems were developed to score persuasive essays and assessed for reliability. Then the agreement of the developed AES systems with human raters was assessed through correlational analyses. We used extrinsic convergent validation approaches to examine if the human and LLM scores correlated with linguistic components. Results indicate strong reliability and agreement for the LLM scores. In terms of convergent validity, initial correlational analyses indicated that the representation LLM AES system showed differential correlations with the human scores in terms of a text length and type-token ratio component. This result contrasts with the correlational results from the generative LLM AES model, which indicated no differences in associations between the model and human scores with regards to the linguistic components.