formative-assessment
Filtering by topic formative-assessment(5)Clear all filters
- PaperAssessment & Evaluation in Higher Education8 Jul 2026
Artificial intelligence and feedback in university education: effectiveness and student perceptions
Valentina Grion, Beatrice Doria, Daniele Agostini, Giorgia Slaviero
A quasi-experimental study with 238 students compared AI-generated feedback from GPT-o4-mini and DeepSeek R1 against expert human feedback in a project-based university course. All feedback conditions led to significant and equivalent improvements in project performance, with no differences in student perceptions. The findings indicate that the pedagogical context, rather than the feedback source, determines effectiveness, supporting AI's role in formative assessment when strong assessment literacy is present.
Original abstract
The integration of generative artificial intelligence (AI) into Higher Education has intensified debates about the role of technology in formative assessment. This study examines the effectiveness and practical comparability of AI-generated feedback in a project-based university course, comparing two large language models (GPT-o4-mini and DeepSeek R1) with feedback provided by an expert human teacher. Adopting a quasi-experimental design, 47 student groups (N = 238) were randomly assigned to one of three feedback conditions. Changes in project performance were analysed using non-parametric tests, robust models, and non-inferiority and equivalence analyses. Students’ perceptions were also assessed through a validated questionnaire (N = 200). Results showed significant improvement in project performance from pre- to post-feedback across all conditions (rrb = 0.77), with no significant differences between feedback sources. Equivalence analyses indicated practical comparability between GPT-o4-mini and teacher feedback, while DeepSeek R1 demonstrated non-inferiority. Students’ perceptions of mastery, emotions, and satisfaction were similarly high across conditions. Findings suggest that feedback effectiveness depends less on its source than on the pedagogical architecture in which it is embedded. When supported by strong assessment literacy and explicit criteria, AI-generated feedback can function as a credible component of formative assessment in higher education.
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of ChatGPT-powered chatbot interactions in a low-stakes formative speaking assessment. Comparing a 290,000-word chatbot corpus with the British National Corpus 2014 revealed that chatbot output resembled written language more than spoken, with higher lexical density and fewer spoken features like stance markers. The framework is designed to be adaptable to future AI conversational agents.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.
- PaperAssessing Writing5 Jun 2026
Beyond the grade: Reimagining second language (L2) writing assessment through ungrading
Shiyu Tang, Shulin Yu
This paper explores ungrading as an alternative to traditional grading in second language writing assessment, arguing that it can promote more meaningful feedback and reduce anxiety. It reimagines assessment as a learning-oriented process rather than a summative judgment.
- PaperAssessment & Evaluation in Higher Education29 May 2026
When technological momentum overshadows pedagogical alignment: a systematic review of AI-generated formative feedback in higher education
Ezgi Çallı, Erkan Er
This systematic review of 103 studies (2020–2025) finds that AI-generated formative feedback in higher education has expanded rapidly, but often lacks deep pedagogical alignment. Positive learner perceptions and efficiency gains are common, yet feedback quality evaluations remain indirect and short-term. The review identifies a structural alignment challenge across pedagogical theory, system design, and practice, arguing that long-term value depends on principled coordination rather than technical sophistication alone.
Original abstract
Providing timely and pedagogically meaningful formative feedback remains a persistent challenge in higher education. Advances in generative artificial intelligence (AI), particularly large language models (LLMs), have accelerated research on automating and augmenting feedback processes. This systematic review synthesises 103 empirical studies published between 2020 and 2025 to examine how AI-generated formative feedback is conceptualised, implemented, and evaluated in higher education. The analysis reveals rapid technological expansion, with AI most commonly positioned as a supplementary assistant to enhance feedback efficiency and scalability. While studies frequently report positive learner perceptions and improvements in feedback-related outcomes, evaluations of feedback quality are often indirect and grounded primarily in short-cycle interventions and perceptual measures. Theoretical grounding is uneven and instructor involvement often remains supervisory. Drawing on these patterns, the review identifies a structural alignment challenge across three interdependent layers: foundational pedagogical theory, system design, and interactional practice. The findings suggest that the long-term educational value of AI-generated formative feedback depends less on technical sophistication alone than on principled coordination between pedagogical intent, technological architecture, and human-AI collaboration. The review clarifies structural patterns in the literature and outlines priorities for theory-informed and context-sensitive implementation.
- PaperBritish Journal of Educational Technology4 May 2026
Student profiles of change in formative assessment behaviour: Replication and evaluation for grade prediction
Oleksandra Poquet, Jelena Jovanovic, Stephan Krusche
This study replicates a complex dynamical systems approach to analyze changes in formative assessment submission patterns among 1362 programming students. It identifies three student profiles of behavioural change, finding that higher entropy in recurrence patterns correlates with better performance and timeliness. While these dynamics-based features do not surpass conventional metrics in prediction, they offer complementary insights for student interventions.
Original abstract
As students learn and practice new skills in university courses, their behaviour can change in response to competing demands and increasing content complexity. However, most metrics used to evaluate study behaviour focus on the number or sequence of activities rather than on the change of behaviour. To address this, we replicate and extend a complex dynamical systems approach to characterise recurrence in behavioural patterns and whether it changes.Using assessment logs from 1362 students in the first 5 weeks of a semester‐long programming course, we examine whether changes in the patterns of formative assessment submissions can differentiate student sub‐groups and predict their performance. We identify three student profiles of behavioural change. We find that higher entropy of recurrence in assessment submission patterns is associated with better performance, and that changes in this entropy signal upcoming changes in performance. We also show that higher entropy of recurrence is associated with greater timeliness of submissions. Finally, we evaluate the predictive value of early behavioural patterns and find that while student profiles of change do not outperform conventional predictive metrics, they offer complementary insights that can enable timely interpretations of student data and inform interventions. Overall, our findings extend the generalisability of behavioural metrics based on complex dynamical systems by demonstrating consistent patterns across courses, LMS types and data sources. Practitioner notes What is already known about the topic Recurrence quantification analysis can capture dynamics of student behaviour. Prior work proposed a methodology based on recurrence of behaviour in a complex system to quantify study behaviour with trace data. Students whose behavioural patterns showed consistently high entropy of recurrence performed better. What this paper adds This study replicates a CDS‐based methodology in a new context, with a typical data source: traces of student assessment submissions. Dynamics‐based features are associated with the timeliness of student submissions and course performance. Dynamics‐based features do not outperform conventional LA features in predicting student performance, but offer complementary insights. Implications for policy/practice More research is needed to interpret what dynamics‐based features mean for teaching practice before they can be acted on. Future research and teaching activities could integrate interviews and self‐reported instruments to examine potential interpretations of dynamics‐based features, such as students' propensity to adapt.