qualitative-coding
Filtering by topic qualitative-coding(1)Clear all filters
- PaperComputers and Education: Artificial Intelligence17 Jul 2026
Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
Kamila Misiejuk, Sonsoles López-Pernas, Eduardo A. Oliveira, Brendan Eagan et al.
The study compares human and LLM ordered coding of qualitative text data from learners, finding systematic differences in structural, transitional, and code-level metrics. Two evaluation approaches are proposed to assess coding quality, emphasizing that LLM errors could cascade and distort automated feedback.
Original abstract
Automating the process of qualitatively coding text data from learners has been a long-standing ambition of learning analytics researchers since it represents an essential step toward delivering timely and scalable feedback. Automating this process is especially challenging in the case of ordered coding schemes —necessary for temporal analytical methods— where one text utterance can be assigned more than one qualitative code and the assignment order matters. This problem goes beyond multi-class and multi-label classification and, therefore, cannot be easily tackled using classic language models such as BERT. Recent advances in generative artificial intelligence, especially with the advent of large language models, have —allegedly— created a substantial step forward in making the goal of automatically coding complex temporal data attainable. However, little is yet known about how to implement this process in a way that most closely resembles human coding, i.e., taking into account the context in which the textual data appears for accurate interpretation. Moreover, due to the complexity of the data and its shape, the accuracy of the results cannot be computed using classic accuracy metrics. This study makes two main contributions: first, it presents two evaluation approaches for assessing the quality of ordered data coding and the usability of LLM in automatically coding ordered processes; and second, it demonstrates a method of LLM prompting that leverages a consistent context window. Our results reveal systematic and statistically significant differences between LLM and human coding across structural, transitional, and code-level metrics for binary and ordered tasks. As classification errors can propagate through automated feedback systems, relying on LLM outputs risks amplifying inaccuracies and producing misleading interpretations of learning processes.