formative-assessment
Filtering by topic formative-assessment(5)Clear all filters
- PaperAssessment & Evaluation in Higher Education8 Jul 2026
Artificial intelligence and feedback in university education: effectiveness and student perceptions
Valentina Grion, Beatrice Doria, Daniele Agostini, Giorgia Slaviero
A quasi-experimental study compared AI-generated feedback from two large language models (GPT-o4-mini and DeepSeek R1) with expert human feedback in a university project-based course. Results showed significant improvement in performance across all conditions, with no practical differences between AI and human feedback, suggesting that pedagogical design matters more than the source of feedback. Students' perceptions of mastery, emotions, and satisfaction were similarly high regardless of feedback source.
Original abstract
The integration of generative artificial intelligence (AI) into Higher Education has intensified debates about the role of technology in formative assessment. This study examines the effectiveness and practical comparability of AI-generated feedback in a project-based university course, comparing two large language models (GPT-o4-mini and DeepSeek R1) with feedback provided by an expert human teacher. Adopting a quasi-experimental design, 47 student groups (N = 238) were randomly assigned to one of three feedback conditions. Changes in project performance were analysed using non-parametric tests, robust models, and non-inferiority and equivalence analyses. Students’ perceptions were also assessed through a validated questionnaire (N = 200). Results showed significant improvement in project performance from pre- to post-feedback across all conditions (rrb = 0.77), with no significant differences between feedback sources. Equivalence analyses indicated practical comparability between GPT-o4-mini and teacher feedback, while DeepSeek R1 demonstrated non-inferiority. Students’ perceptions of mastery, emotions, and satisfaction were similarly high across conditions. Findings suggest that feedback effectiveness depends less on its source than on the pedagogical architecture in which it is embedded. When supported by strong assessment literacy and explicit criteria, AI-generated feedback can function as a credible component of formative assessment in higher education.
- PaperDOAJ — Language assessment1 Jun 2026
The Effects of E-Portfolio Implementation Based on Formative Experimental Design on Students’ English Reading and Writing Skills and Self-Efficacy
Özlem Ören, Azmi Türkan
A formative experimental design with 29 tenth graders showed that e-portfolio implementation improved English reading and writing achievement scores and self-efficacy beliefs. Qualitative data indicated students perceived e-portfolios as beneficial for learning and skill development.
Original abstract
The present study aimed to examine the use of e-portfolios in relation to students’ English reading and writing skills and self-efficacy beliefs. Because e-portfolio is believed to have a potential to foster 21st century skills including creativity and digital literacy. The participants consisted of 29 10th graders in a state high school. Formative experimental design was employed a research design involving a single experimental group. In the quantitative dimension, the reading and writing subscales of the "English Self-Efficacy Belief Scale" and "English Reading and Writing Achievement Test," developed by the researcher, were used in order to collect the data. So as to score English writing achievement test, "The Newly Developed Essay Scoring Rubric" was utilized to assess students’ writing performance. For the qualitative part of the research, a semi-structured interview form prepared by the researcher and researcher notes were used to collect the data. Quantitative data were analyzed using dependent-samples t-tests to compare pre-test and post-test scores. Qualitative data were analyzed through content analysis. The findings of the study revealed improvements in participants’ English reading and writing achievement scores as well as their self-efficacy beliefs following the e-portfolio implementation. Qualitative findings also suggested that students perceived e-portfolios as beneficial for supporting their learning processes and developing their language skills.
- PaperAssessment & Evaluation in Higher Education29 May 2026
When technological momentum overshadows pedagogical alignment: a systematic review of AI-generated formative feedback in higher education
Ezgi Çallı, Erkan Er
A systematic review of 103 empirical studies on AI-generated formative feedback in higher education reveals rapid technological expansion, with AI commonly positioned as a supplementary assistant. Evaluations of feedback quality are often based on short-cycle interventions and perceptual measures, lacking deep pedagogical grounding. The review identifies a structural alignment challenge across foundational theory, system design, and practice, emphasizing that long-term value depends on principled coordination between pedagogy, technology, and human-AI collaboration.
Original abstract
Providing timely and pedagogically meaningful formative feedback remains a persistent challenge in higher education. Advances in generative artificial intelligence (AI), particularly large language models (LLMs), have accelerated research on automating and augmenting feedback processes. This systematic review synthesises 103 empirical studies published between 2020 and 2025 to examine how AI-generated formative feedback is conceptualised, implemented, and evaluated in higher education. The analysis reveals rapid technological expansion, with AI most commonly positioned as a supplementary assistant to enhance feedback efficiency and scalability. While studies frequently report positive learner perceptions and improvements in feedback-related outcomes, evaluations of feedback quality are often indirect and grounded primarily in short-cycle interventions and perceptual measures. Theoretical grounding is uneven and instructor involvement often remains supervisory. Drawing on these patterns, the review identifies a structural alignment challenge across three interdependent layers: foundational pedagogical theory, system design, and interactional practice. The findings suggest that the long-term educational value of AI-generated formative feedback depends less on technical sophistication alone than on principled coordination between pedagogical intent, technological architecture, and human-AI collaboration. The review clarifies structural patterns in the literature and outlines priorities for theory-informed and context-sensitive implementation.
- PaperJournal of Learning Analytics18 May 2026
Learning-Aware Reliability Estimation for Tutor Skill Assessment Using Large Language Models
Conrad Borchers, Danielle R. Thomas, Jionghao Lin, Kenneth R. Koedinger
A novel Rasch-based split-half method adjusts reliability estimates for learning gains in pre-post assessments scored by large language models. Applied to GPT-4 scoring of 985 tutors' open-ended responses, the method achieved satisfactory reliability (0.733) with as few as 14 items, and identified three skill subdimensions: socio-emotional, cognitive, and fairness-related tutoring skills.
Original abstract
Assessment is foundational to learning analytics, especially in evaluating instructional interventions and guiding improvement in online learning environments. With the growing use of large language models (LLMs) to score open-ended responses, questions arise about the reliability of these model-generated scores, particularly in short pre-post formats where learners are expected to improve. This study introduces a novel method for estimating test reliability that adjusts for learning gains using a Rasch-based split-half approach. We validated this approach through simulation under realistic conditions of missing data and score change, showing tangible improvements in reliability estimation compared to baseline methods. Applying this method to a dataset of 985 tutors completing 12 online lessons, we find that GPT-4-based scoring achieves satisfactory reliability, with open-ended responses (0.733) outperforming multiple-choice items (0.652). Both item types jointly yielded the highest reliability (0.774). Hence, as few as 14 open-ended items (across an average of 3-4 completed lessons) were sufficient to surpass common reliability thresholds of 0.7 or higher. Principal component analysis revealed a skill structure with a strong primary dimension shared across almost all lessons and interpretable subdimensions—socio-emotional, cognitive, and fairness-related tutoring skills—supporting a bifactor-like model. These findings demonstrate that GPT-4 and similar LLMs can be effectively used for formative assessment of complex instructional skills in online and personalized learning contexts, provided their reliability is empirically verified. This study contributes an open-source, learning-aware framework for scalable and reliable AI-supported assessment in learning analytics contexts.
- PaperBritish Journal of Educational Technology4 May 2026
Student profiles of change in formative assessment behaviour: Replication and evaluation for grade prediction
Oleksandra Poquet, Jelena Jovanovic, Stephan Krusche
Replicating a complex dynamical systems approach with assessment logs from 1362 students in a programming course, the study identifies three profiles of change in formative assessment submission behavior. Higher entropy of recurrence in submission patterns was associated with better performance and timeliness, offering complementary insights for grade prediction beyond conventional metrics.
Original abstract
As students learn and practice new skills in university courses, their behaviour can change in response to competing demands and increasing content complexity. However, most metrics used to evaluate study behaviour focus on the number or sequence of activities rather than on the change of behaviour. To address this, we replicate and extend a complex dynamical systems approach to characterise recurrence in behavioural patterns and whether it changes.Using assessment logs from 1362 students in the first 5 weeks of a semester‐long programming course, we examine whether changes in the patterns of formative assessment submissions can differentiate student sub‐groups and predict their performance. We identify three student profiles of behavioural change. We find that higher entropy of recurrence in assessment submission patterns is associated with better performance, and that changes in this entropy signal upcoming changes in performance. We also show that higher entropy of recurrence is associated with greater timeliness of submissions. Finally, we evaluate the predictive value of early behavioural patterns and find that while student profiles of change do not outperform conventional predictive metrics, they offer complementary insights that can enable timely interpretations of student data and inform interventions. Overall, our findings extend the generalisability of behavioural metrics based on complex dynamical systems by demonstrating consistent patterns across courses, LMS types and data sources. Practitioner notes What is already known about the topic Recurrence quantification analysis can capture dynamics of student behaviour. Prior work proposed a methodology based on recurrence of behaviour in a complex system to quantify study behaviour with trace data. Students whose behavioural patterns showed consistently high entropy of recurrence performed better. What this paper adds This study replicates a CDS‐based methodology in a new context, with a typical data source: traces of student assessment submissions. Dynamics‐based features are associated with the timeliness of student submissions and course performance. Dynamics‐based features do not outperform conventional LA features in predicting student performance, but offer complementary insights. Implications for policy/practice More research is needed to interpret what dynamics‐based features mean for teaching practice before they can be acted on. Future research and teaching activities could integrate interviews and self‐reported instruments to examine potential interpretations of dynamics‐based features, such as students' propensity to adapt.