chatgpt
Filtering by topic chatgpt(5)Clear all filters
- PaperComputers & Education24 Jun 2026
Color me confounded: A critical analysis of media comparisons on ChatGPT in education
Alyssa P. Lawson, Amédee Marchand Martella, Joshua Weidlich, Miriam Mulders et al.
A critical analysis examines how media comparisons of ChatGPT in education often lead to confusion and misrepresentation of the tool's capabilities.
- PaperAssessing Writing23 Jun 2026
“Investigating the impact of ChatGPT-assisted self-assessment on college students’ writing development: Insights from diverse linguistic backgrounds”
Hamidreza Moeiniasl
Examined how ChatGPT-assisted self-assessment affects college students' writing development, with a focus on learners from diverse linguistic backgrounds.
- PaperReCALL5 May 2026
Involving AI in the development of intercultural competence: Exploring the role of ChatGPT in fostering cognitive flexibility
Isabelle Drewelow
This study investigates whether interacting with ChatGPT supports perspective shifting and cognitive flexibility in a university-level LSP French course. Analysis of student interviews reveals that while concerns about AI content legitimacy existed, the tool enabled some students to adopt emic stances and reflect on interpretive limits. Prompt literacy emerged as crucial for fostering dynamic partnerships with ChatGPT for intercultural competence development.
Original abstract
This study investigates whether and how interacting with ChatGPT may offer a context that supports perspective shifting and the development of cognitive flexibility, defined as the capacity to move between etic (outsider) and emic (insider) perspectives. Drawing on individual interviews with students enrolled in an advanced university-level Language for Specific Purposes (LSP) French course focused on marketing and advertising in France, this qualitative study examines students’ perspectives on their experiences using ChatGPT to conduct market research on French consumer needs and preferences. The analysis reveals that while students expressed concerns about the legitimacy, authenticity, and cultural positioning of AI-generated content, the interactive and conversational nature of the tool enabled some students to experiment with culturally unfamiliar roles, adopt emerging emic stances, and reflect on the limits of their interpretive frameworks. However, co-creative engagement or shared agency with ChatGPT was not automatic and depended on prompt design, tolerance for ambiguity, and the negotiation of subjective positioning. Rather than facilitating perspective transformation, ChatGPT-supported interactions appeared to foster more modest but meaningful shifts in interpretive positioning and dialectical thinking. The study points to prompt literacy as crucial for fostering more dynamic partnerships with ChatGPT and enabling students to explore alternative perspectives and roles in ways that support the development of intercultural competence in the L2 classroom.
- PaperLanguage Testing25 Mar 2026
™ChatGPT for automated writing evaluation: Scoring and feedback across prompt conditions
Yewon Lee, Myunghwan Hwang
Six ChatGPT models with different prompt configurations were tested against human raters on 60 EFL writing samples. Prompt design significantly influenced scoring consistency and severity, with Chain-of-Thought and Fill-in-the-blank prompts yielding higher reliability. Learners perceived the feedback positively, but reasoning-intensive domains still required human oversight.
Original abstract
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
- PaperReCALL29 Dec 2025
Impact of prompt sophistication on ChatGPT’s output for automated written corrective feedback
Na Luo, Yifan Wang, Zhe (Victor) Zhang, Yile Zhou et al.
Compared ChatGPT's automated written corrective feedback (AWCF) accuracy across generic and domain-specific prompts against Grammarly. Found that domain-specific prompts, especially one-shot, significantly improved error detection, with zero-shot matching Grammarly and one-shot surpassing it. However, even the most sophisticated prompt still showed limitations compared to Grammarly.
Original abstract
The emergence of large language models, exemplified by ChatGPT, has garnered growing attention for their potential to generate feedback in second language writing, particularly automated written corrective feedback (AWCF). In this study, we examined how prompt design – a generic prompt and two domain-specific prompts (zero-shot and one-shot) enriched with comprehensive domain knowledge about written corrective feedback (WCF) – influences ChatGPT’s ability to provide AWCF. The accuracy and coverage of ChatGPT’s feedback across these three prompts were benchmarked against Grammarly, a widely used traditional automated writing evaluation (AWE) tool. We find that ChatGPT’s ability in flagging language errors grew considerably with prompt sophistication driven by the integration of domain-specific knowledge and examples. While the generic prompt resulted in substantially lower performance than Grammarly, the zero-shot prompt achieved comparable results to it and the one-shot prompt surpassed it considerably in error detection. Notably, the most pronounced improvement in ChatGPT’s performance was observed in its detection of frequent error categories, including those of word choice or expression, direct translation, sentence structure and pronoun. Nonetheless, even with the most sophisticated prompt, ChatGPT still displayed certain limitations when compared to Grammarly. Our study has both theoretical and practical implications. Theoretically, it lends empirical evidence to Knoth et al .’s (2024) proposition to separate domain-specific AI literacy from generic AI literacy. Practically, it sheds light on the pedagogical application and technical development of AWE systems.