learning-analytics
Filtering by topic learning-analytics(9)Clear all filters
- PaperComputers and Education: Artificial Intelligence17 Jul 2026
Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
Kamila Misiejuk, Sonsoles López-Pernas, Eduardo A. Oliveira, Brendan Eagan et al.
This study compares human and large language model (LLM) ordered coding of qualitative learner data, revealing systematic differences that cascade through temporal analysis. Two evaluation approaches assess coding quality and demonstrate a context-aware prompting method. Results show significant disparities in structural, transitional, and code-level metrics, cautioning against relying on LLM outputs for automated feedback in learning analytics.
Original abstract
Automating the process of qualitatively coding text data from learners has been a long-standing ambition of learning analytics researchers since it represents an essential step toward delivering timely and scalable feedback. Automating this process is especially challenging in the case of ordered coding schemes —necessary for temporal analytical methods— where one text utterance can be assigned more than one qualitative code and the assignment order matters. This problem goes beyond multi-class and multi-label classification and, therefore, cannot be easily tackled using classic language models such as BERT. Recent advances in generative artificial intelligence, especially with the advent of large language models, have —allegedly— created a substantial step forward in making the goal of automatically coding complex temporal data attainable. However, little is yet known about how to implement this process in a way that most closely resembles human coding, i.e., taking into account the context in which the textual data appears for accurate interpretation. Moreover, due to the complexity of the data and its shape, the accuracy of the results cannot be computed using classic accuracy metrics. This study makes two main contributions: first, it presents two evaluation approaches for assessing the quality of ordered data coding and the usability of LLM in automatically coding ordered processes; and second, it demonstrates a method of LLM prompting that leverages a consistent context window. Our results reveal systematic and statistically significant differences between LLM and human coding across structural, transitional, and code-level metrics for binary and ordered tasks. As classification errors can propagate through automated feedback systems, relying on LLM outputs risks amplifying inaccuracies and producing misleading interpretations of learning processes.
- PaperJournal of Learning Analytics8 Jul 2026
Toward Reliable Estimation of Algorithmic Bias for Minority Groups
Jaeyoon Choi, Shamya Karumbaiah, Jeffrey Matayoshi
Predictive models in learning analytics often show performance disparities across demographic groups, but reliable estimation of group bias is hampered by small group sizes and sampling error, especially for marginalized students. Using simulations and real-world data, the study recommends bootstrapping for confidence intervals, using multiple metrics, and moving beyond p-values to improve bias estimation.
Original abstract
While predictive models are widely used in learning analytics, several studies have shown that the performance of these models can vary significantly across different demographic groups of students. The first step to audit for and mitigate these group biases is to accurately estimate them. However, the current practices for identifying and measuring group bias often suffer from reliability issues. In this paper, we use simulations and real-world data analysis to explore statistical factors that impact the reliable estimate of group bias and suggest approaches to improve their statistical robustness. Our analysis revealed that small group sizes lead to high variability in group bias estimation due to sampling error -- an issue that is more likely to impact students from historically marginalized communities. We then suggest statistical approaches, such as bootstrapping, to construct confidence intervals for a more reliable estimation of group bias. Based on our findings, we encourage future learning analytics researchers to ensure sufficiently large group sizes, construct confidence intervals rather than relying on p-values, use at least two metrics, and move beyond the dichotomy of the presence or absence of bias for a more comprehensive evaluation of group bias.
- PaperBritish Journal of Educational Technology3 Jul 2026
A balance between stability and flexibility: Adaptive patterns of self‐regulated learning processes shape game‐based learning
Elizabeth B. Cloude, Stefan E. Huber, Jingwei Wei, Bianca Esmanhoto et al.
This study analyzes self-regulated learning (SRL) as a complex system in game-based learning (GBL), using multimodal data to examine patterns and interactions among cognitive, affective, metacognitive, and motivational processes. Results show a quadratic relationship between the regularity of cognitive, affective, and metacognitive state transitions and learning outcomes, indicating an optimal balance between stability and flexibility. Physiological measures of motivation did not predict learning performance, suggesting motivation's role may be mediated through cognitive and metacognitive engagement.
Original abstract
To ensure learning efficiency in game‐based learning (GBL), learners must regulate cognitive, affective, metacognitive and motivational (CAMM) processes, collectively known as self‐regulated learning (SRL). SRL is dynamic and non‐linear, characterized by regulatory patterns and CAMM interactions that lead to macro‐level SRL behaviours. In this paper, we explore SRL as a complex system by analysing patterns and interactions among CAMM processes during GBL and examine their relation to learning outcomes. Thirty‐seven ( n = 37) healthy adults used Antidote COVID‐19, a GBL environment designed to increase emotional engagement and biology knowledge. To define momentary CAMM processes, subjective measures of cognitive, affective and metacognitive states via think‐ and emote‐alouds were synchronized with two physiological correlates of motivation: intensity of facial expressions (arousal) and skin conductance response. Sequence analysis and recurrence quantification defined the regularity of patterns within CAMM components, and transfer entropy estimated concurrent interactions between CAMM components. Our results indicated a quadratic relationship between the turbulence (a measure of regularity) of cognitive, affective and metacognitive (CAM) state transitions (via think‐ and emote‐aloud methods) provided the best fit, explaining 31% of variability in learning. This highlights an optimal balance between stability and flexibility in CAM state transitions that were beneficial for learning. However, physiological correlates of motivation were not predictive regarding learning performance. This lack of predictive capability of the considered measures may reflect their limitations in capturing nuanced motivation dynamics relevant to learning or suggest that motivation's role in learning is predominantly mediated through cognitive and metacognitive engagement rather than directly influencing outcomes. Practitioner notes What is already known about this topic Game‐based learning (GBL) is a promising intervention for improving science knowledge, but it places high self‐regulatory demands on learners. Learners must manage cognitive, affective, motivational and metacognitive (CAMM) processes to ensure GBL efficiency; otherwise known as self‐regulated learning (SRL). Prior research has shown that CAMM processes are each important for learning in games, but they are often studied in isolation rather than as interacting processes over time. What this paper adds This study conceptualizes SRL in GBL as a complex system utilizing mixed multimodal data, emphasizing patterns and interactions among CAMM processes rather than static averages. Results show that learning is maximized when learners exhibit an optimal balance between stability and flexibility in their cognitive, affective and metacognitive state transitions, neither overly rigid nor overly random regulation patterns. Cognitive, affective and metacognitive processes (assessed via think–/emote‐alouds) were strongly related to learning gain, whereas physiological indicators of motivation alone were not predictive. Implications for practice and/or policy Designers of GBL environments should support adaptive regulation, encouraging learners to flexibly shift SRL strategies and emotions while maintaining coherence in their learning process. Educators and researchers should be cautious about relying solely on physiological measures (eg, arousal) for assessing learning effectiveness, as these may not capture meaningful regulatory dynamics. Policies and evaluation frameworks for educational games should prioritize tools and analytics that capture process‐level SRL patterns over time, rather than focusing exclusively on outcomes or isolated behavioural indicators.
- PaperJournal of Learning Analytics21 Jun 2026
Practitioner-Informed Learning Analytics Metrics for Measuring Curricular Complexity for Transfer Students
David Reeping, Dustin Grote, Julie Christensen, Sulabh Khadka et al.
Introduced three new metrics—inflexibility factor, transfer delay factor, and credit loss—to better measure curricular complexity for transfer students. Validated these metrics through focus groups with 38 transfer professionals, who affirmed their alignment with real-world curricular barriers.
Original abstract
This study expands the concept of curricular complexity as defined in the Curricular Analytics framework (Heileman et al., 2018) by providing qualitative evidence of the suitability of three new metrics, concerning timing of course offerings, extended time-to-degree, and credit loss, that more adequately address curricular challenges encountered by transfer students. Curricular Analytics is a method for analyzing a curriculum that enables practitioners and researchers to quantify and systematically analyze the impacts of course sequencing in a plan of study on student outcomes. However, the original conceptualization falls short of capturing the substantive challenges faced by transfer students who enter an undergraduate program at various points in the curricular sequence. This study was guided by the following research question: “How do three new measures of curricular complexity (i.e., inflexibility factor, transfer delay factor, and credit loss) align with transfer professionals’ perceptions of curricular barriers for transfer students?” Using a grounded theory approach, we conducted seven focus groups with 38 transfer professionals across the United States. We presented these transfer experts with each new measure and prompted them to reflect on its validity based on their experiences supporting transfer students. We found transfer professionals resonated strongly with all three new metrics, suggesting strong initial construct and content validity.
- PaperComputers and Education: Artificial Intelligence19 Jun 2026
Federated and explainable learning analytics for privacy-preserving academic risk modeling across heterogeneous educational institutions
William Villegas-Ch, Alexandra Maldonado Navarro, Jaime Govea, Joselin García-Ortiz et al.
This study proposes a federated, explainable learning analytics framework for modeling academic risk across heterogeneous educational institutions, integrating temporal behavioral features with socio-academic data under a multitask learning scheme. Results show that federated models maintain discriminative performance and stable convergence across non-IID partitions, while calibration metrics are more sensitive to distributional shifts. Explainability analysis reveals that feature importance remains structurally stable across institutions, supporting the deployment of privacy-preserving models in diverse settings.
Original abstract
The increasing digitization of higher education has enabled the development of learning analytics models to identify students at risk of academic failure or dropout; however, most existing approaches rely on centralized training and assume homogeneous data distributions, limiting their applicability across institutions with heterogeneous student populations and interaction patterns, while privacy constraints restrict data sharing and hinder collaborative model development. To address these challenges, this study proposes a federated, explainable learning analytics framework for modeling academic risk trajectories under controlled institutional heterogeneity. The proposed architecture integrates temporal behavioral representations with socio-academic features within a multitask learning scheme, evaluated under both centralized and federated regimes, while modeling institutional heterogeneity through parameterized non-IID partitions that introduce controlled class imbalance, temporal drift, and feature-level variability. Experimental results show that federated models preserve strong discriminative performance and stable convergence as heterogeneity increases. At the same time, calibration metrics exhibit greater sensitivity to distributional shifts, revealing a decoupling between ranking performance and probabilistic reliability. In parallel, explainability analysis shows that the relative importance of features remains structurally stable across institutions, despite variations in contribution magnitudes. Cross-platform evaluation further shows that models retain discriminative capacity when transferred across educational environments, while exhibiting changes in calibration and explanatory intensity. These findings highlight the importance of multidimensional evaluation in federated learning systems, jointly considering performance, calibration, and interpretability, and provide a methodological framework for deploying robust and privacy-preserving learning analytics models in heterogeneous educational settings.
- PaperAssessment & Evaluation in Higher Education4 Jun 2026
University instructors’ contemporary assessment literacy: development and validation of a questionnaire
Goudarz Alibakhshi
Developed and validated a 35-item questionnaire measuring university instructors' contemporary assessment literacy, covering nine dimensions including AI-responsive assessment, digital assessment literacy, and learning analytics. The instrument showed satisfactory reliability and validity based on expert reviews and factor analyses with 462 instructors. Findings highlight the need for measures that capture emerging assessment competencies in higher education.
Original abstract
Assessment literacy has become a key professional competence in higher education, where instructors are expected to design learning-oriented, ethical, inclusive, digitally mediated and evidence-informed assessment practices. However, existing measures do not fully capture contemporary demands related to feedback, learner involvement, artificial intelligence, learning analytics, accessibility and assessment consequences. This study developed and validated a questionnaire measuring university instructors’ contemporary assessment literacy. Using a multiphase, mixed-methods instrument development design, the study was conducted in two sequential phases. In Phase 1, a preliminary 39-item pool was reviewed by 22 assessment experts from three universities in Tehran. Expert ratings supported item relevance, clarity, representativeness and essentiality, and 12 items were revised. In Phase 2, the revised questionnaire was administered to 670 university instructors from four Tehran universities; 462 usable responses were returned. Exploratory factor analysis supported a nine-factor solution explaining 66.66% of the variance. The retained dimensions were learning-oriented assessment literacy, feedback literacy, learner involvement, fairness and ethics, digital assessment literacy, AI-responsive assessment literacy, inclusive and accessible assessment literacy, consequential validity and washback literacy, and assessment data and learning analytics literacy. Confirmatory factor analysis supported the final 35-item model, with satisfactory reliability, convergent validity, discriminant validity and model fit.
- PaperJournal of Learning Analytics18 May 2026
Learning-Aware Reliability Estimation for Tutor Skill Assessment Using Large Language Models
Conrad Borchers, Danielle R. Thomas, Jionghao Lin, Kenneth R. Koedinger
Introduces a learning-aware reliability estimation method using a Rasch-based split-half approach to adjust for learning gains when assessing LLM scoring reliability. Results show GPT-4 scoring achieves satisfactory reliability, with open-ended items (0.733) outperforming multiple-choice (0.652), and both combined yielding the highest reliability (0.774). The findings support using LLMs for formative assessment of complex instructional skills in online learning contexts.
Original abstract
Assessment is foundational to learning analytics, especially in evaluating instructional interventions and guiding improvement in online learning environments. With the growing use of large language models (LLMs) to score open-ended responses, questions arise about the reliability of these model-generated scores, particularly in short pre-post formats where learners are expected to improve. This study introduces a novel method for estimating test reliability that adjusts for learning gains using a Rasch-based split-half approach. We validated this approach through simulation under realistic conditions of missing data and score change, showing tangible improvements in reliability estimation compared to baseline methods. Applying this method to a dataset of 985 tutors completing 12 online lessons, we find that GPT-4-based scoring achieves satisfactory reliability, with open-ended responses (0.733) outperforming multiple-choice items (0.652). Both item types jointly yielded the highest reliability (0.774). Hence, as few as 14 open-ended items (across an average of 3-4 completed lessons) were sufficient to surpass common reliability thresholds of 0.7 or higher. Principal component analysis revealed a skill structure with a strong primary dimension shared across almost all lessons and interpretable subdimensions—socio-emotional, cognitive, and fairness-related tutoring skills—supporting a bifactor-like model. These findings demonstrate that GPT-4 and similar LLMs can be effectively used for formative assessment of complex instructional skills in online and personalized learning contexts, provided their reliability is empirically verified. This study contributes an open-source, learning-aware framework for scalable and reliable AI-supported assessment in learning analytics contexts.
- PaperComputer Assisted Language Learning5 May 2026
Learning analytics on multimodal GAI-driven EFL oral learning: uncovering learning behavior clusters with motivation and performance dynamics
Yuting Chen, Morris Siu-Yung Jong, Michael Yi-Chao Jiang, Ming Li
This study uses learning analytics to analyze multimodal data from generative AI-driven EFL oral learning, identifying clusters of learning behaviors and examining their relationships with motivation and performance dynamics.
- PaperBritish Journal of Educational Technology4 May 2026
Student profiles of change in formative assessment behaviour: Replication and evaluation for grade prediction
Oleksandra Poquet, Jelena Jovanovic, Stephan Krusche
Using assessment logs from 1362 students in a programming course, this study replicates a complex dynamical systems approach to characterize changes in formative assessment submission patterns. Three student profiles of behavioural change were identified, with higher entropy of recurrence associated with better performance and timeliness. While these dynamics-based features do not outperform conventional metrics for grade prediction, they offer complementary insights for interpreting student data.
Original abstract
As students learn and practice new skills in university courses, their behaviour can change in response to competing demands and increasing content complexity. However, most metrics used to evaluate study behaviour focus on the number or sequence of activities rather than on the change of behaviour. To address this, we replicate and extend a complex dynamical systems approach to characterise recurrence in behavioural patterns and whether it changes.Using assessment logs from 1362 students in the first 5 weeks of a semester‐long programming course, we examine whether changes in the patterns of formative assessment submissions can differentiate student sub‐groups and predict their performance. We identify three student profiles of behavioural change. We find that higher entropy of recurrence in assessment submission patterns is associated with better performance, and that changes in this entropy signal upcoming changes in performance. We also show that higher entropy of recurrence is associated with greater timeliness of submissions. Finally, we evaluate the predictive value of early behavioural patterns and find that while student profiles of change do not outperform conventional predictive metrics, they offer complementary insights that can enable timely interpretations of student data and inform interventions. Overall, our findings extend the generalisability of behavioural metrics based on complex dynamical systems by demonstrating consistent patterns across courses, LMS types and data sources. Practitioner notes What is already known about the topic Recurrence quantification analysis can capture dynamics of student behaviour. Prior work proposed a methodology based on recurrence of behaviour in a complex system to quantify study behaviour with trace data. Students whose behavioural patterns showed consistently high entropy of recurrence performed better. What this paper adds This study replicates a CDS‐based methodology in a new context, with a typical data source: traces of student assessment submissions. Dynamics‐based features are associated with the timeliness of student submissions and course performance. Dynamics‐based features do not outperform conventional LA features in predicting student performance, but offer complementary insights. Implications for policy/practice More research is needed to interpret what dynamics‐based features mean for teaching practice before they can be acted on. Future research and teaching activities could integrate interviews and self‐reported instruments to examine potential interpretations of dynamics‐based features, such as students' propensity to adapt.