ai-feedback
Filtering by topic ai-feedback(8)Clear all filters
- PaperComputers & Education1 Jul 2026
Developing L2 writing student self-assessment literacy through AI-generated feedback and AI chain-of-thought: An action research study with lower-proficiency university EFL learners
Pan Ye, Shulin Yu, Icy Lee, Chenggang Liang
An action research study investigates how AI-generated feedback and AI chain-of-thought promote self-assessment literacy in L2 writing among lower-proficiency university EFL learners.
- PaperAssessment & Evaluation in Higher Education23 Jun 2026
AI in the gatekeeper’s chair: elite researchers’ perceptions of AI-assisted feedback in journal peer review
Heng Li
Elite researchers (Nature/Science authors) perceive AI-assisted peer reviews as less fair, useful, and acceptable than human reviews, and show aversion to both the AI itself and researchers who delegate reviews to AI. This suggests AI integration may erode trust in peer review unless human oversight is preserved.
Original abstract
Peer review serves as the structural foundation of scientific integrity, yet the system currently faces unprecedented strain. In response, AI tools are being deployed to assist human expert review, a development that has generated considerable debate. However, the psychological impact of this transition on the scientific community remains only partially understood. Our study utilises a sequential mixed-methods approach to examine the perceptions of AI-mediated feedback among a cohort of elite researchers (Nature and Science authors). Quantitative findings from our randomised experimental survey (N = 495) demonstrated that AI-assisted reviews were viewed as deficient in fairness, usefulness, and acceptance compared to human-led evaluations. Qualitative evidence from 47 in-depth interviews further identified a dual-layered aversion to AI-assisted feedback, concerning both the technology (AI use) and the agent (AI user). Specifically, scholars perceived AI-generated critiques as devoid of the necessary disciplinary nuance required for high-stakes evaluation. Moreover, a notable ‘AI user aversion’ emerged: reviewers who delegated tasks to AI were perceived as lacking the diligence and empathic engagement essential to the peer-review contract. Together, these finding suggest that the integration of AI into peer review may erode trust in the research evaluation process and journals should implement robust governance that preserves human oversight.
- PaperJournal of Second Language Writing20 Jun 2026
Question-only AI Socratic dialogue as dialogic feedback in L2 argumentative writing: A quasi-experimental study
Li-Jen Wang
The paper reports a quasi-experimental study investigating the use of question-only AI Socratic dialogue as a form of dialogic feedback for L2 argumentative writing.
- PaperAssessing Writing8 Jun 2026
Examining linguistic reasoning and metalinguistic strategies in L2 writing: Insights from teacher- and AI-feedback revisions
Huican Huo, Lawrence Jun Zhang
A longitudinal study of 10 L2 writers over 12 weeks found that repeated feedback-revision cycles with teacher and AI suggestions shifted learners from correctness-focused edits to deeper metalinguistic reasoning, including structural comparison, self-explanation, and integrative argumentation. Composite reasoning scores increased, with largest gains in depth of explanation, and patterns emerged linking feedback characteristics to specific reasoning behaviors.
Original abstract
Second language (L2) writers’ capacity to reason about language choices during revision is theoretically central to advanced writing development, yet few studies trace how this linguistic reasoning emerges in repeated feedback-revision cycles that integrate teacher and AI suggestions. This multi-case, longitudinal mixed-methods study examined how ten L2 learners enacted metalinguistic strategies and developed linguistic reasoning across a 12-week online Academic English program. Triangulated data sources included 1246 feedback revision episodes, pre-post reasoning-task responses, and semi-structured exit interviews. We operationalized linguistic reasoning as the processual construct of interest and treated metalinguistic strategies as observable indicators; reasoning responses were scored with a three-dimensional analytic rubric while revisions and interaction logs were coded thematically and by strategy. Results show a systematic shift from early, correctness-focused edits toward later revisions characterized by coordinated structural comparison, more explicit self-explanation, and integrative argumentation; composite reasoning scores increased, with the largest gains observed on depth of explanation. Process analyses identified recurrent patterns in which alternative-rich feedback coincided with comparison moves, partially divergent suggestions often coincided with more explicit self-explanation, and integrative argumentation involving discourse-level revision became more visible over time. Pedagogical implications are also discussed.
- PaperAssessing Writing3 Jun 2026
Using GenRewrite to provide personalized feedback for form-function alignment in EAP writing
Yizhe Wei, Yue Wang, Tan Jin
The study introduces GenRewrite, a tool that provides personalized feedback to help EAP writers align form and function in their writing.
- PaperAssessment & Evaluation in Higher Education1 Jun 2026
The human touch of feedback: students’ experiences of CARE in peer versus AI-generated feedback
Lan Li, Jiming Zhou
A study comparing AI and peer feedback in an interpreting course found that while AI offered comprehensive, criterion-based comments with a positive tone, peer feedback demonstrated greater developmental sensitivity and relational grounding. Students valued the contextual understanding and authentic support from peers, leading to the CARE framework (Care respect, Attainable goals, Relational recognition, Emphasized problem identification). The findings suggest that productive AI integration should design complementary systems rather than simulating human touch.
Original abstract
The integration of AI-generated feedback into higher education has increased feedback volume and efficiency. Yet concerns persist that it lacks the ‘human touch’, a construct that remains undertheorised and empirically unexamined. To examine what constitutes the human touch, this study compared AI and peer feedback in an interpreting course, capturing 41 university students’ immediate responses through the think-aloud method across seven weeks. Analysis revealed that whereas AI provided comprehensive, criterion-based commentary with a more positive tone, peer feedback demonstrated greater developmental sensitivity and relational grounding. Students showed emotional indifference to AI feedback but valued the contextual understanding and authentic support that peer feedback provided. These patterns informed the empirically grounded CARE framework: Care respect, Attainable goals, Relational recognition, and Emphasised problem identification. Each CARE dimension depends on qualities emerging from shared participation in learning communities that algorithmic systems struggle to replicate. Theoretically, CARE offers concrete dimensions for understanding feedback effectiveness beyond content coverage. Productive AI integration requires not simulating human touch but designing complementary systems that leverage the strengths of different feedback sources. The presence of human feedback providers does not guarantee the human touch, either. The CARE dimensions demand deliberate assessment design and invite further exploration.
- PaperReCALL2 Mar 2026
Evaluating GPT-generated feedback for beginning-level Spanish writing: An exploratory study
Alyssia Miller De Rutté, Maite Correa
A custom GPT model, 'Belinda,' was trained to score and provide feedback on A1-level Spanish writing. Its scoring showed moderate alignment with human raters but fell short of reliability benchmarks, and its feedback was often vague, incomplete, or inaccurate. The study concludes that the model is not yet a viable substitute for human evaluation and emphasizes the need for human-AI collaboration.
Original abstract
Feedback is integral to second language (L2) writing instruction. However, large class sizes and limited teacher time often challenge the delivery of personalized feedback, prompting interest in AI-powered solutions such as ChatGPT (Escalante et al., 2023; Huete-García & Tarp, 2024; Steiss et al., 2024; Yoon et al., 2023; Zhang, 2024). This study evaluates a task-customized GPT model, “Belinda,” trained to assess A1-level Spanish learners’ writing and provide feedback. Two research questions guided the investigation: (1) Can Belinda accurately score beginner Spanish writing using a provided rubric? (2) Can Belinda deliver constructive qualitative feedback? Human and GPT-generated scores were compared for inter- and intrarater reliability, and qualitative analyses categorized the feedback for usability in the classroom. Results revealed moderate alignment between Belinda’s scores and human raters, though reliability of the GPT fell short of calibration benchmarks. Feedback quality varied, with Belinda often providing vague, incomplete, or inaccurate suggestions. Despite iterative training, the GPT struggled to balance error correction with encouragement, a critical need for novice learners. Additionally, inconsistencies in identical GPT versions raised concerns about reliability. While Belinda showed potential in automating feedback, its limitations in accuracy, contextual understanding, and positivity suggest it is not yet a viable substitute for human evaluation by itself. These findings emphasize the challenges of integrating AI into L2 instruction and call for the need for extensive datasets, robust training, and human–AI collaboration to achieve pedagogically sound outcomes. Future research should explore hybrid feedback models and scalable solutions to enhance AI’s role in language education without compromising learner progress or confidence.
- PaperERIC — Assessment & second language1 Jan 2025
Feedforwarding Diagnostic Language Assessment: Artificial Intelligence- (AI-) Driven Weakness Identification and Contextualised Feedback for Second Language Speaking
Shungo Suzuki, Hiroaki Takatsu, Ryuki Matsuura, Miina Koyama et al.
The study proposes a new diagnostic language assessment (DLA) approach for speaking that uses AI to identify lexical weaknesses and provides contextualized feedback integrated with remedial activities. In an experiment with 59 Japanese learners of English, the group receiving diagnostic feedback outperformed the control group on a posttest and maintained learning gains, while the control group showed only temporary improvement from task repetition alone.
Original abstract
The current study proposes a new approach to weakness identification in diagnostic language assessment (DLA) for speaking skills. We also propose to design actionable and contextualised diagnostic feedback through the systematic integration of feedback and remedial learning activities. Focusing on lexical use in second language speaking, the current study developed and validated our DLA programme in terms of actual learning gains, using an experimental design. A total of 59 beginner-to-intermediate-level Japanese learners of English were randomly assigned to control or experimental groups. While both groups engaged in task repetition with a conversational artificial intelligence (AI) agent on six occasions, only the experimental group received the diagnostic feedback on lexical use including the paraphrased utterances of their original utterance. The results showed that the control group (task repetition only) demonstrated significant improvement during the task repetition sessions but failed to transfer and retain the learning gains. In contrast, despite the lack of practice effects, the experimental group (task repetition with diagnostic feedback) outperformed the control group at the posttest with a near-medium effect size. A qualitative investigation into learners' perceptions further confirmed that the proposed contextualised diagnostic feedback succeeded in heightening their awareness of weaknesses.