second-language-writing
Filtering by topic second-language-writing(11)Clear all filters
- PaperLanguage Teaching24 Jul 2026
Writing assessment literacy and feedback literacy in second language writing
Zhicheng Mao, Shulin Yu, Icy Lee
This paper explores the intersection of writing assessment literacy and feedback literacy in the context of second language writing education.
- PaperETS Research Report Series16 Jul 2026
Second Language Writing Processes in the TOEFL iBT® Test: Examining the Write for an Academic Discussion Task Using Eye-Tracking, Keystroke Logging, and Stimulated Recall
Ching-Ni Hsieh, Renka Ohta
Eye-tracking and keystroke logging reveal that the TOEFL iBT Write for an Academic Discussion task engages L2 writers in cognitive processes aligning with theoretical models of academic writing, supporting construct validity. Frequent referencing of task materials and substantial revision behaviors were observed, with five core cognitive processes identified: prompt (re)reading, planning, formulation, monitoring, and revision.
Original abstract
The present study investigates the construct validity of the TOEFL iBT® writing test and focuses on the writing processes elicited by the Write for an Academic Discussion (WAD) task. Nineteen adult L2 English users completed the WAD task while their eye movements and keystrokes were recorded, followed by stimulated recall interviews to capture their writing strategies and thought processes. Eye-tracking data revealed sustained attention to the writing window, with frequent reference to the task question and example posts. Keystroke logging indicated substantial initial pauses and language-related revision behaviors. Qualitative analysis of stimulated recalls identified five core cognitive processes: prompt (re)reading, planning, formulation, monitoring, and revision. The integration of eye-tracking, keystroke logging, and stimulated recall demonstrates that the WAD task engages L2 writers in processes consistent with theoretical models of academic writing, supplying backing for the construct validity of the TOEFL iBT Writing test. The findings also reveal that task design features may influence L2 writing processes and shape writers’ strategic behaviors and attention allocation. Suggested citation: Hsieh, C.-N., & Ohta, R. (in press). Second language writing processes in the TOEFL iBT® test: Examining the Write for an Academic Discussion task using eye-tracking, keystroke logging, and stimulated recall (Research Report). ETS. https://doi.org/10.64634/tk1sd682
- PaperJournal of Second Language Writing16 Jul 2026
Conceptualizing voice in the age of AI: A response to Sandstead and Kibler (2025)
Paul Kei Matsuda, Xiao Tan
This response engages with Sandstead and Kibler (2025) to reconceptualize the notion of voice in writing within the context of AI technologies, addressing implications for L2 writing pedagogy and assessment.
- PaperComputers & Education14 Jul 2026
LLM-derived metrics in second language writing assessment: an explainable AI approach
Jingying Hu, Yan Cong
This study derives interpretable metrics from LLM internal representations—surprisal, perplexity, and embedding similarity—to assess second language writing. Applied to 1,196 Chinese L2 learner essays, these metrics correlate with proficiency and complement traditional linguistic features, improving classification accuracy. The findings support transparent, reproducible AI-assisted writing assessment for low-resource languages.
Original abstract
Large language models (LLMs) are increasingly used in educational technology for automated writing assessment, yet most applications rely on prompt-based scoring and feedback generation, which often lack transparency, reproducibility, and interpretability. This study investigates whether model-internal LLM representations can provide interpretable and reproducible metrics for second language (L2) writing assessment. We derived surprisal and perplexity from next-token prediction to quantify linguistic predictability and embedding-based similarity to measure semantic coherence across sentences. These metrics were computed at the token, sentence, and discourse levels using three pretrained Chinese language models and evaluated on 1,196 essays written by Chinese L2 learners across four proficiency levels. Their relationships with 11 established linguistic measures of fluency, lexical sophistication, phraseological complexity, and syntactic complexity were also examined. Results showed that surprisal and perplexity generally decreased with proficiency for the two Traditional Chinese-focused models, indicating greater linguistic predictability in more proficient writing, whereas the multilingual model showed weaker sensitivity. Embedding-based similarity increased with proficiency, reflecting stronger semantic coherence. Combining LLM-derived metrics with classical linguistic features improved proficiency classification and prediction beyond either feature set alone. Correlation and qualitative analyses further demonstrated that the proposed metrics capture complementary aspects of writing while revealing conditions under which their interpretations become less reliable. These findings demonstrate the value of interpretable, model-derived metrics for transparent, reproducible, and scalable AI-supported L2 writing assessment, particularly for underrepresented learner populations and lower-resource languages.
- PaperAssessing Writing5 Jun 2026
Beyond the grade: Reimagining second language (L2) writing assessment through ungrading
Shiyu Tang, Shulin Yu
Proposes reimagining second language writing assessment by moving beyond traditional grading to 'ungrading' approaches.
- PaperJournal of Second Language Writing2 Jun 2026
A good relationship is not enough: How teacher trust shapes feedback-seeking behaviors in second language writing
Mostafa Papi, Jeannine Turner, Wenting Song, Hadya Soliman et al.
This study investigates how teacher trust influences students' feedback-seeking behaviors in second language writing, finding that trust is a crucial factor beyond a good relationship.
- PaperJournal of Second Language Writing1 Jun 2026
JSLW Editorial for June 2026
Stephen Doolan, Mimi Li
An editorial for the June 2026 issue of the Journal of Second Language Writing is presented, likely introducing the issue's contents and themes relevant to second language writing research and pedagogy.
- PaperJournal of Second Language Writing15 May 2026
Non-empirical scholarship in the Journal of Second Language Writing
Xueyi Yuan
This paper examines the role and characteristics of non-empirical scholarship published in the Journal of Second Language Writing.
- PaperComputer Assisted Language Learning4 May 2026
Feedback timing and engagement with feedback. Effects on L2 written accuracy
Florentina Nicolás-Conesa, Lourdes Cerezo, Sophie McBride
Investigates how timing of feedback and learners' engagement with it affect second language written accuracy.
- PaperComputer Assisted Language Learning29 Apr 2026
Bing-powered revision: an exploratory study of EFL students’ use of Bing for L2 writing revision
Yingmin Wang, Chun Lai, Manfei Xu, Tan Jin
An exploratory study investigated how EFL students utilized the Bing search engine for revising their L2 writing.
- PaperReCALL29 Dec 2025
Impact of prompt sophistication on ChatGPT’s output for automated written corrective feedback
Na Luo, Yifan Wang, Zhe (Victor) Zhang, Yile Zhou et al.
Compared ChatGPT's automated written corrective feedback (AWCF) accuracy across generic and domain-specific prompts against Grammarly. Found that domain-specific prompts, especially one-shot, significantly improved error detection, with zero-shot matching Grammarly and one-shot surpassing it. However, even the most sophisticated prompt still showed limitations compared to Grammarly.
Original abstract
The emergence of large language models, exemplified by ChatGPT, has garnered growing attention for their potential to generate feedback in second language writing, particularly automated written corrective feedback (AWCF). In this study, we examined how prompt design – a generic prompt and two domain-specific prompts (zero-shot and one-shot) enriched with comprehensive domain knowledge about written corrective feedback (WCF) – influences ChatGPT’s ability to provide AWCF. The accuracy and coverage of ChatGPT’s feedback across these three prompts were benchmarked against Grammarly, a widely used traditional automated writing evaluation (AWE) tool. We find that ChatGPT’s ability in flagging language errors grew considerably with prompt sophistication driven by the integration of domain-specific knowledge and examples. While the generic prompt resulted in substantially lower performance than Grammarly, the zero-shot prompt achieved comparable results to it and the one-shot prompt surpassed it considerably in error detection. Notably, the most pronounced improvement in ChatGPT’s performance was observed in its detection of frequent error categories, including those of word choice or expression, direct translation, sentence structure and pronoun. Nonetheless, even with the most sophisticated prompt, ChatGPT still displayed certain limitations when compared to Grammarly. Our study has both theoretical and practical implications. Theoretically, it lends empirical evidence to Knoth et al .’s (2024) proposition to separate domain-specific AI literacy from generic AI literacy. Practically, it sheds light on the pedagogical application and technical development of AWE systems.