learner-corpus
Filtering by topic learner-corpus(1)Clear all filters
- PaperComputers & Education14 Jul 2026
LLM-derived metrics in second language writing assessment: an explainable AI approach
Jingying Hu, Yan Cong
This study derived surprisal, perplexity, and embedding-based similarity from large language models to quantify linguistic predictability and semantic coherence in second language writing. These metrics, computed on Chinese learner essays, decreased with proficiency for Traditional Chinese-focused models, while embedding similarity increased. Combining LLM-derived metrics with classical linguistic features improved proficiency classification, supporting transparent and scalable AI assessment for underrepresented languages.
Original abstract
Large language models (LLMs) are increasingly used in educational technology for automated writing assessment, yet most applications rely on prompt-based scoring and feedback generation, which often lack transparency, reproducibility, and interpretability. This study investigates whether model-internal LLM representations can provide interpretable and reproducible metrics for second language (L2) writing assessment. We derived surprisal and perplexity from next-token prediction to quantify linguistic predictability and embedding-based similarity to measure semantic coherence across sentences. These metrics were computed at the token, sentence, and discourse levels using three pretrained Chinese language models and evaluated on 1,196 essays written by Chinese L2 learners across four proficiency levels. Their relationships with 11 established linguistic measures of fluency, lexical sophistication, phraseological complexity, and syntactic complexity were also examined. Results showed that surprisal and perplexity generally decreased with proficiency for the two Traditional Chinese-focused models, indicating greater linguistic predictability in more proficient writing, whereas the multilingual model showed weaker sensitivity. Embedding-based similarity increased with proficiency, reflecting stronger semantic coherence. Combining LLM-derived metrics with classical linguistic features improved proficiency classification and prediction beyond either feature set alone. Correlation and qualitative analyses further demonstrated that the proposed metrics capture complementary aspects of writing while revealing conditions under which their interpretations become less reliable. These findings demonstrate the value of interpretable, model-derived metrics for transparent, reproducible, and scalable AI-supported L2 writing assessment, particularly for underrepresented learner populations and lower-resource languages.