formative-assessment
Filtering by topic formative-assessment(2)Clear all filters
- PaperAssessment & Evaluation in Higher Education8 Jul 2026
Artificial intelligence and feedback in university education: effectiveness and student perceptions
Valentina Grion, Beatrice Doria, Daniele Agostini, Giorgia Slaviero
A quasi-experimental study with 238 students compared AI-generated feedback from GPT-o4-mini and DeepSeek R1 against expert human feedback in a project-based university course. All feedback conditions led to significant and equivalent improvements in project performance, with no differences in student perceptions. The findings indicate that the pedagogical context, rather than the feedback source, determines effectiveness, supporting AI's role in formative assessment when strong assessment literacy is present.
Original abstract
The integration of generative artificial intelligence (AI) into Higher Education has intensified debates about the role of technology in formative assessment. This study examines the effectiveness and practical comparability of AI-generated feedback in a project-based university course, comparing two large language models (GPT-o4-mini and DeepSeek R1) with feedback provided by an expert human teacher. Adopting a quasi-experimental design, 47 student groups (N = 238) were randomly assigned to one of three feedback conditions. Changes in project performance were analysed using non-parametric tests, robust models, and non-inferiority and equivalence analyses. Students’ perceptions were also assessed through a validated questionnaire (N = 200). Results showed significant improvement in project performance from pre- to post-feedback across all conditions (rrb = 0.77), with no significant differences between feedback sources. Equivalence analyses indicated practical comparability between GPT-o4-mini and teacher feedback, while DeepSeek R1 demonstrated non-inferiority. Students’ perceptions of mastery, emotions, and satisfaction were similarly high across conditions. Findings suggest that feedback effectiveness depends less on its source than on the pedagogical architecture in which it is embedded. When supported by strong assessment literacy and explicit criteria, AI-generated feedback can function as a credible component of formative assessment in higher education.
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of ChatGPT-powered chatbot interactions in a low-stakes formative speaking assessment. Comparing a 290,000-word chatbot corpus with the British National Corpus 2014 revealed that chatbot output resembled written language more than spoken, with higher lexical density and fewer spoken features like stance markers. The framework is designed to be adaptable to future AI conversational agents.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.