second-language-writing
Filtering by topic second-language-writing(4)Clear all filters
- PaperLanguage Teaching24 Jul 2026
Writing assessment literacy and feedback literacy in second language writing
Zhicheng Mao, Shulin Yu, Icy Lee
This paper explores the intersection of writing assessment literacy and feedback literacy in the context of second language writing education.
- PaperETS Research Report Series16 Jul 2026
Second Language Writing Processes in the TOEFL iBT® Test: Examining the Write for an Academic Discussion Task Using Eye-Tracking, Keystroke Logging, and Stimulated Recall
Ching-Ni Hsieh, Renka Ohta
Eye-tracking and keystroke logging reveal that the TOEFL iBT Write for an Academic Discussion task engages L2 writers in cognitive processes aligning with theoretical models of academic writing, supporting construct validity. Frequent referencing of task materials and substantial revision behaviors were observed, with five core cognitive processes identified: prompt (re)reading, planning, formulation, monitoring, and revision.
Original abstract
The present study investigates the construct validity of the TOEFL iBT® writing test and focuses on the writing processes elicited by the Write for an Academic Discussion (WAD) task. Nineteen adult L2 English users completed the WAD task while their eye movements and keystrokes were recorded, followed by stimulated recall interviews to capture their writing strategies and thought processes. Eye-tracking data revealed sustained attention to the writing window, with frequent reference to the task question and example posts. Keystroke logging indicated substantial initial pauses and language-related revision behaviors. Qualitative analysis of stimulated recalls identified five core cognitive processes: prompt (re)reading, planning, formulation, monitoring, and revision. The integration of eye-tracking, keystroke logging, and stimulated recall demonstrates that the WAD task engages L2 writers in processes consistent with theoretical models of academic writing, supplying backing for the construct validity of the TOEFL iBT Writing test. The findings also reveal that task design features may influence L2 writing processes and shape writers’ strategic behaviors and attention allocation. Suggested citation: Hsieh, C.-N., & Ohta, R. (in press). Second language writing processes in the TOEFL iBT® test: Examining the Write for an Academic Discussion task using eye-tracking, keystroke logging, and stimulated recall (Research Report). ETS. https://doi.org/10.64634/tk1sd682
- PaperJournal of Second Language Writing16 Jul 2026
Conceptualizing voice in the age of AI: A response to Sandstead and Kibler (2025)
Paul Kei Matsuda, Xiao Tan
This response engages with Sandstead and Kibler (2025) to reconceptualize the notion of voice in writing within the context of AI technologies, addressing implications for L2 writing pedagogy and assessment.
- PaperComputers & Education14 Jul 2026
LLM-derived metrics in second language writing assessment: an explainable AI approach
Jingying Hu, Yan Cong
This study derives interpretable metrics from LLM internal representations—surprisal, perplexity, and embedding similarity—to assess second language writing. Applied to 1,196 Chinese L2 learner essays, these metrics correlate with proficiency and complement traditional linguistic features, improving classification accuracy. The findings support transparent, reproducible AI-assisted writing assessment for low-resource languages.
Original abstract
Large language models (LLMs) are increasingly used in educational technology for automated writing assessment, yet most applications rely on prompt-based scoring and feedback generation, which often lack transparency, reproducibility, and interpretability. This study investigates whether model-internal LLM representations can provide interpretable and reproducible metrics for second language (L2) writing assessment. We derived surprisal and perplexity from next-token prediction to quantify linguistic predictability and embedding-based similarity to measure semantic coherence across sentences. These metrics were computed at the token, sentence, and discourse levels using three pretrained Chinese language models and evaluated on 1,196 essays written by Chinese L2 learners across four proficiency levels. Their relationships with 11 established linguistic measures of fluency, lexical sophistication, phraseological complexity, and syntactic complexity were also examined. Results showed that surprisal and perplexity generally decreased with proficiency for the two Traditional Chinese-focused models, indicating greater linguistic predictability in more proficient writing, whereas the multilingual model showed weaker sensitivity. Embedding-based similarity increased with proficiency, reflecting stronger semantic coherence. Combining LLM-derived metrics with classical linguistic features improved proficiency classification and prediction beyond either feature set alone. Correlation and qualitative analyses further demonstrated that the proposed metrics capture complementary aspects of writing while revealing conditions under which their interpretations become less reliable. These findings demonstrate the value of interpretable, model-derived metrics for transparent, reproducible, and scalable AI-supported L2 writing assessment, particularly for underrepresented learner populations and lower-resource languages.