Topics
(655)Clear all filters
- PaperApplied Linguistics16 Jul 2026
The role of teacher gaze in speaking turns during interactive book reading: a multiple case mobile-eye-tracking study
Thibaut Duthois, Susanna Isotalo, Ymke Taillieu, Ruben Vanderlinde et al.
Using mobile eye-tracking, this study examined how teacher gaze shapes turn-taking during interactive book reading in early childhood education. Results showed unequal distribution of teacher attention, with teachers' sustained visual fixations strongly associated with both teacher and child speaking turns, indicating teacher-driven participation and interactional inequalities.
Original abstract
Gaze is foregrounded as a central interactional resource positioning individuals as speakers or bystanders. In early childhood education (ECE), where direct involvement in interactions with the teacher is central to language development and where unequal participation may lay the groundwork for far-reaching processes of educational inequality, research into gaze remains scarce. This study, including six teachers and thirty children, examines variation in children’s involvement during book reading activities and analyses how teachers shape these conversational dynamics. Using epistemic network analysis alongside multimodal quantitative and qualitative conversation analyses, this study provides a fine-grained account of multimodal turn-taking. The findings reveal unequal distribution of teacher attention, resulting in unequal opportunities for children to participate. Results revealed a relationship between teachers’ fixations (i.e. sustained moments of visual attention) and both teachers’ and children’s speaking turns, underscoring the central importance of gaze in conversational dynamics in classrooms. Conversation analysis further shows that children’s participation is largely teacher-driven, with most child turns being initiated by the teacher. Overall, the results indicate interactional inequalities in ECE.
- PaperApplied Linguistics16 Jul 2026
Research Methods for Applied Linguistics: A Practical Guide
Bahram Kazemian
This practical guide covers key research methods used in applied linguistics, offering step-by-step procedures for conducting studies in language teaching and learning.
- PaperApplied Linguistics16 Jul 2026
AI and the simplification of task design
Anna Mendoza
This commentary critiques the trend of using generative AI in academic writing and qualitative research, arguing that AI simplifies complex tasks. The author contends that AI's standardized representations undermine the nuanced understanding needed in EAP writing and qualitative data analysis. The piece ultimately questions whether AI should be used at all in these contexts.
Original abstract
This commentary begins by summarizing Jeon et al.’s (2025) concerns about Generative Artificial Intelligence (GenAI)’s tendency to standardize language, followed by Hu’s (2026) response, on how to bring Artificial Intelligence (AI) into teaching (academic) writing in ways that allow for negotiation of forms. My own response, given the bigger picture, is whether AI needs to be used at all. First, as a former English for Academic Purposes (EAP) writing instructor, I argue that AI represents writing tasks differently from humans, and that going along with its simplified representation is a step back after human experts have rendered more complex representations of the tasks. Second, as a qualitative researcher, I argue that AI is of limited use in qualitative data analysis, since what qualitative researchers analyze goes beyond what is in the text. I conclude that in both these cases, AI simplifies tasks and detracts from what students of academic writing or novice qualitative researchers need to learn.
- PaperarXiv — AI in Education (cs.CY)15 Jul 2026
When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring
Nischal Ashok Kumar, Payu Wittawatolarn, Sana Kang, Marisa C. Peczuh et al.
The paper investigates cross-rubric generalization in automated essay scoring, where models trained on essays scored under one rubric must perform on new rubrics targeting different aspects. Using a trait-based intermediate representation and target-essay supervision, the approach improves macro F1 by 5% in the hardest setting. Their best open-source Llama-based model outperforms GPT-5-mini prompting by 2.1% and trails GPT-5 by 1.9%.
Original abstract
Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target essays are unseen during training. We further find that increasing target-essay supervision improves performance, with our best fine-tuned open-source Llama-based model outperforming GPT-5-mini prompting by 2.1% macro F1 and trailing GPT-5 by 1.9%. These results show that trait-based intermediate structure and controlled supervision improve generalization to unseen rubrics.
- PaperTESOL Quarterly15 Jul 2026
Effects of Motivational Interventions on EFL Learners' Learning Effort and Achievement: An L2 Motivational Self System Perspective
Xuejun Ye, Guangwei Hu
A quasi-experimental study with 391 Chinese junior secondary EFL students examined the effects of three motivational interventions based on the L2 motivational self system. Traditional, vision-based, and combined strategies all improved learning effort and achievement immediately, but only vision-based and combined interventions sustained gains in achievement over four months.
Original abstract
Grounded in the L2 motivational self system theory, this quasi‐experimental study investigated the impact of three motivational interventions on two outcome measures: L2 learning effort and L2 achievement. A total of 391 Chinese junior secondary students were assigned to four conditions: a control group receiving no motivational instruction, and three experimental groups exposed to either traditional motivational strategies (MSs), vision‐based MSs, or a combined approach integrating both types of strategy. The effects of the interventions were assessed through questionnaires and tests administered at three time points to monitor changes in both outcome variables. The respective and relative effects of the interventions were determined by two‐way mixed design ANOVAs and post‐hoc pairwise comparisons. Results indicated that all three motivational treatments produced significantly positive effects on both outcome variables immediately following the interventions. All three motivational treatments sustained their positive effects on L2 learning effort, but only the vision‐based and combined treatments maintained their impact on L2 achievement. Moreover, all motivational treatments had comparable effects on L2 learning effort both during and 4 months after the interventions. Although all interventions were similarly effective on L2 achievement during the treatment period, the vision‐based and combined interventions outperformed the traditional MS intervention in terms of sustained effects.
- PaperarXiv — AI in Education (cs.CY)14 Jul 2026
Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates
Lynn Vonderhaar, Juan Couder, Siri Siqveland, Omar Ochoa et al.
AI techniques, specifically Large Language Models (LLMs), were applied to analyze and revise curricular patterns in an undergraduate Software Engineering degree. This automated approach reduces the time needed for curriculum changes and helps identify bottlenecks that delay graduation, potentially improving on-time graduation rates.
Original abstract
The rise of Artificial Intelligence (AI) enables automatic analysis of large amounts of data. Previously time-consuming and labor-intensive tasks can be completed much more efficiently with the use of AI. This work uses AI techniques to analyze and revise curricular patterns in an undergraduate degree for Software Engineering. Curricula often have long sequences where failure to pass a class within the sequence may jeopardize completion of the degree within four years. Manual analysis and revision of curricula by university faculty is a lengthy and labor-intensive process, causing changes to occur rarely and making it impossible to keep up with the changing needs of students. This work reduces the time-to-change for curricula and reduces bottlenecks and graduation delays by using Large Language Models (LLMs) to analyze curricular patterns and suggest revisions.
- PaperarXiv — AI in Education (cs.CY)14 Jul 2026
A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential
Gendo Kumoi, Fumie Watanabe, Tota Suko, Takashi Ishida et al.
Large Language Models were used to build a semi-automated system generating dialogue-based Text-to-Speech lessons based on cognitive apprenticeship theory. A study with 245 high school students found that dialogue TTS significantly improved comprehension and cognitive engagement compared to single-speaker TTS, though single TTS was perceived as more natural. The system augments, not replaces, educators through a human-in-the-loop workflow.
Original abstract
This study proposes a semi-automated system for generating dialogue-based lessons using Large Language Models (LLMs) and Text-to-Speech (TTS) technology, and exploratorily examines its educational potential via a practical quasi-experiment. The system augments rather than replaces educators through a three-stage human-in-the-loop workflow (LLM-based slide/narration generation, educator review, automated audiovisual integration), and introduces a novel method for generating Expert-Novice dialogue narration based on cognitive apprenticeship theory. In a study of 245 first-year high school students who sequentially experienced three lesson formats (instructor voice, single-speaker TTS, dialogue TTS; content differed across sessions, limiting format/content separation), we conducted within-subject (Friedman test, N<=183) and repeated cross-sectional (Mann-Whitney U, N=229/206) analyses. TTS audio did not substantially degrade the learning experience versus instructor voice, supported by TOST equivalence testing. Dialogue TTS was significantly superior to single TTS in comprehension (p=.006, q=.025) and cognitive engagement (p=.019, q=.048); enjoyment was non-significant after FDR correction (q=.081) but reached significance after controlling for prior knowledge (proportional-odds model, OR=1.65, q=.025), and these advantages were not attributable to prior-knowledge imbalance. Conversely, single TTS was superior in audio naturalness (p<.001, q<.001, r=-.238), revealing a trade-off between dialogue's benefits and higher extraneous cognitive load. Dialogue format was preferred by 66.9% of learners as most enjoyable (p<.001). These results reflect a fixed-order design; replication is needed before generalizing them as effects of lesson format. This study provides a theoretical and empirical basis for the educational acceptability of TTS audio and for TTS lesson-format design.
- PaperAssessment & Evaluation in Higher Education14 Jul 2026
Evaluating the effectiveness of an awareness-based intervention on reducing accent bias in students’ evaluations of instructors
Esther Sosa-Herrera, Kat Silaj, Mary J. Keushkerian, Melissa Paquette-Smith
A brief awareness-based intervention showed mixed results in reducing accent bias in student evaluations of instructors. In two experiments, students consistently rated a Mandarin-accented instructor more negatively than an American-accented instructor, despite equal learning outcomes. The intervention reduced bias in the first experiment but was less effective with different materials and a more diverse sample.
Original abstract
Student evaluations of teaching can be biased by the ‘way’ the instructor speaks or their accent. In the current study, we investigate whether a brief awareness-based intervention can be effective in reducing accent-based bias in students’ evaluations of instructors. In a series of controlled laboratory experiments, college student participants were randomly assigned to watch a short lecture narrated by either an instructor who spoke English with a Mandarin accent or an American accent. Prior to evaluating the instructor, participants in the intervention condition were given instructions designed to reduce bias, whereas participants in the control condition proceeded directly to the evaluation form. Although there were no statistically significant differences in performance on a multiple-choice assessment of learning, participants in both experiments showed evidence of bias, rating the Mandarin-accented instructor more negatively than the American-accented instructor. In Experiment 1, the intervention reduced the disparity in participants’ overall ratings of the two instructors; however, in Experiment 2, the intervention was less effective with different lecture materials and a more diverse sample of students. Awareness-based interventions may show promise in reducing real-world biases in college classes and could support efforts to improve diversity and inclusivity in higher education.
- PaperComputers and Education: Artificial Intelligence14 Jul 2026
Balancing AI responsibility with privacy, safety, and utility: Unlearning in large language models for mathematics education
Chenglu Li, Gökhan Gülfidan, Yinqi Zhang-Kopf
This study applies gradient-based LLM unlearning to reduce personally identifiable information (PII) and harmful content in math tutoring models while maintaining performance on math tasks. Results show substantial decreases in PII and harm rates without sacrificing utility, demonstrating a path toward responsible AI in education.
Original abstract
Online mathematics learning platforms are increasingly adopting large language models (LLMs) to provide scalable, on-demand support, but these models may reproduce private information from training data or generate harmful language. This raises concerns about responsibility in educational settings regarding the use of pre-trained models. LLM unlearning is an emerging area for reducing a model’s ability to produce specific unwanted content and remains underexplored in educational research. This study aims to investigate how LLM unlearning reduces the model's reliance on personally identifiable information (PII) and inappropriate content in the math tutoring context, while maintaining the model's utility on both single-label and multi-label downstream math tasks. We applied a gradient-based LLM unlearning approach to three different models, which were pre-trained on approximately 3 million data points from an Algebra I online discussion forum between students and professional tutors. PII and harmful content were detected on this training data and used for unlearning in two different orders (PII and harmful content unlearning). Then, the generated outputs from these two unlearning models were compared with those of the pre-trained model in terms of PII-containing output rate and harmful rate. Moreover, unlearned models were evaluated on two different math classification tasks. The results showed that the rates of PII-containing output rate and harmfulness substantially decreased compared to the pre-trained models, and the utility of the unlearned model was still maintained. These findings demonstrate how LLM unlearning can be applied to pre-trained models to behave them more responsibly, while maintaining strong model performance on math-related tasks.
- PaperComputers & Education14 Jul 2026
LLM-derived metrics in second language writing assessment: an explainable AI approach
Jingying Hu, Yan Cong
This study derived surprisal, perplexity, and embedding-based similarity from large language models to quantify linguistic predictability and semantic coherence in second language writing. These metrics, computed on Chinese learner essays, decreased with proficiency for Traditional Chinese-focused models, while embedding similarity increased. Combining LLM-derived metrics with classical linguistic features improved proficiency classification, supporting transparent and scalable AI assessment for underrepresented languages.
Original abstract
Large language models (LLMs) are increasingly used in educational technology for automated writing assessment, yet most applications rely on prompt-based scoring and feedback generation, which often lack transparency, reproducibility, and interpretability. This study investigates whether model-internal LLM representations can provide interpretable and reproducible metrics for second language (L2) writing assessment. We derived surprisal and perplexity from next-token prediction to quantify linguistic predictability and embedding-based similarity to measure semantic coherence across sentences. These metrics were computed at the token, sentence, and discourse levels using three pretrained Chinese language models and evaluated on 1,196 essays written by Chinese L2 learners across four proficiency levels. Their relationships with 11 established linguistic measures of fluency, lexical sophistication, phraseological complexity, and syntactic complexity were also examined. Results showed that surprisal and perplexity generally decreased with proficiency for the two Traditional Chinese-focused models, indicating greater linguistic predictability in more proficient writing, whereas the multilingual model showed weaker sensitivity. Embedding-based similarity increased with proficiency, reflecting stronger semantic coherence. Combining LLM-derived metrics with classical linguistic features improved proficiency classification and prediction beyond either feature set alone. Correlation and qualitative analyses further demonstrated that the proposed metrics capture complementary aspects of writing while revealing conditions under which their interpretations become less reliable. These findings demonstrate the value of interpretable, model-derived metrics for transparent, reproducible, and scalable AI-supported L2 writing assessment, particularly for underrepresented learner populations and lower-resource languages.
- PaperAssessing Writing14 Jul 2026
Can citations indicate performance levels? An investigation of integrated argumentative writing tests
Qin Xie, Chang Zhang
This study investigates whether citations in integrated argumentative writing tests can serve as indicators of different performance levels.
- PaperAssessing Writing14 Jul 2026
Development and evaluation of a student feedback agency scale
Jiaxian Ye, Lawrence Jun Zhang, Helen R. Dixon
A Student Feedback Agency Scale (SFAS) was developed and validated in two phases using Chinese international postgraduate students. Principal Component Analysis yielded six components (Action Taking, Goal Setting, Processing, Generating, Self-efficacy, Seeking), and Confirmatory Factor Analysis confirmed good model fit and structural invariance across gender, academic level, and discipline. The instrument provides a psychometrically sound tool for measuring student agency in writing feedback processes.
Original abstract
Feedback agency is a key concept in enhancing student writing performance. While growing attention has been paid to student feedback agency, existing research remains largely theoretical and qualitative. As a result, there is a lack of psychometrically supported instruments to measure this construct. To address this gap, the present two-phase study aimed to develop and evaluate a Student Feedback Agency Scale (SFAS), drawing on social cognitive theory. In the development stage, Principal Component Analysis (PCA) was conducted on a sample of 235 Chinese international postgraduate students. The results yielded a 23-item SFAS comprising six components: Action Taking, Goal Setting, Processing, Generating, Self-efficacy, and Seeking. Using an independent sample of 349 participants from the same population, Confirmatory Factor Analysis (CFA) was conducted in the evaluation stage. The results supported a good model fit (RMSEA = .055, IFI = .923, TLI = .908, and CFI = .922). Multi-group CFAs further confirmed the structural invariance across gender, academic level, and discipline. Overall, the findings provide psychometric evidence to support the interpretation and use of the SFAS scores to measure student agency in writing feedback processes. Based on these results, the factor structure and subscales of the SFAS are discussed, and implications are outlined.
- PaperarXiv — Language & NLP (cs.CL)13 Jul 2026
A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol
Esteban U. Vega Barajas
A validated protocol for classifying teaching-evaluation comments by category and sentiment was tested for durability and cross-language transfer. Using Spanish and English corpora, the study compared sparse lexical features, frozen transformer embeddings, and prompted LLMs, finding the protocol durable: a 2026 frontier model achieved the highest thematic F1 on Spanish but showed no sentiment advantage over cheaper models, and English sentiment performance was descriptively similar across models. Results indicate that model choice for this task is a deployment decision rather than a property of the protocol.
Original abstract
Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. We re-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.
- PaperarXiv — Language & NLP (cs.CL)13 Jul 2026
Globally Consistent Coloring Schemes for Language Identification
Moses Charikar, Jon Kleinberg, Chirag Pabbaraju
The paper shows that a single terminal bit per string is sufficient to identify any countable collection of infinite languages in Gold's language identification model. The authors construct a global two-color terminal coloring using transfinite recursion, but prove that no Borel-definable finite coloring can achieve the same result.
Original abstract
We study how little extra information is needed to make adversarial language learning possible. In Gold's model of language identification in the limit, a learner is given an enumeration of the strings from an unknown language chosen from a countable language collection. The learner guesses the identity of the language over the course of the enumeration, and it succeeds if, eventually, all of its guesses are the correct language. Classical results of Gold and Angluin show that many natural collections cannot be learned in this way. Recent work on trace colorings, motivated by the success of thinking-trace strategies in language learning, overcomes this obstruction by annotating every symbol of every string with a color. We ask whether the learner really needs this whole sequence of colors, or whether one color at the end of each string (a terminal coloring) is enough for language identification. We show that just one terminal bit per string is enough for every countable collection of infinite languages. In fact, the colorings can be chosen collection-independently: there is a single assignment of a two-color terminal coloring to every infinite language such that the same preassigned colorings identify every countable subcollection. Thus, in this model, an entire color trace can be compressed to one bit attached to the end of each example. Our global construction uses transfinite recursion, and we prove that this kind of nonconstructivity is unavoidable for any bounded number of colors. As a notion of constructivity, we use the formalism of Borel maps (a regularity condition satisfied by natural explicit constructions); we show that no global terminal coloring with a finite number of colors defined by a Borel map can identify all countable subcollections. By contrast, known trace-coloring constructions are Borel when encoded as terminal colorings, but require infinitely many colors.
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis
Adrian-Marius Dumitran, Iulia-Maria Popescu
A 15-nation comparative analysis finds that many students complete secondary education without formal programming exposure, and among those who do, a 'Syntax Ceiling' limits algorithmic depth: Python is widespread but C++ remains in elite STEM tracks. Governance structures and high-stakes exams, rather than curriculum content alone, drive these inequities, undermining the goal of universal AI literacy.
Original abstract
The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by governance decisions made largely with reference to the specialist one. This paper presents a comparative analysis of curricula and examination frameworks across fifteen countries, identifying two structural challenges. First, in several systems a significant portion of students completes secondary education without any formal programming exposure. Second, among those who do receive CS education, a \emph{Syntax Ceiling} emerges: Python-based instruction reaches most students, while the algorithmic depth associated with C++ remains concentrated in elite STEM tracks. Drawing on reform cases spanning centralised mandates (France, China, Japan), assessment-driven systems (Poland, Romania, South Korea), and recent universal reforms (Switzerland, Kazakhstan), we show that governance structures and high-stakes examinations are the primary drivers of both challenges -- and that specialist and general-track language choices are rarely independent, linked through shared teacher pipelines that curriculum policy seldom acknowledges. Achieving genuine AI literacy for all requires confronting not just curriculum content, but the access architectures and resource constraints that determine who receives it -- and at what depth.
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
Alexis Popovici, Andrei Ionascu, Adrian-Marius Dumitran
A systematic audit of four LLMs acting as history tutors found that safety-aligned models exhibit epistemic paternalism, differentially refusing 76.7% of educational requests from low-tier students and reducing access to complex geopolitical content for marginalized learners. The study identifies patterns including differential refusal, epistemic gatekeeping, agency theft, and elite hermeneutics, arguing that current safety alignment functions as a paternalistic filter that perpetuates narrative segregation.
Original abstract
As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regarding the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. We uncover four interconnected patterns of \emph{epistemic paternalism}: (1)~\textbf{Differential Refusal}, where safety-aligned models block 76.7\% of educational requests from low-tier students; (2)~\textbf{Epistemic Gatekeeping}, evidenced by a 3$\times$ reduction in access to geopolitical complexity (e.g., the contested ``coup theory'') for marginalized learners; (3)~\textbf{Agency Theft}, a lexical shift where models like LLaMA produce a 5$\times$ higher victimization-to-politics vocabulary ratio for Roma students compared to elite peers; and (4)~\textbf{Elite Hermeneutics}, where AI tutors disproportionately withhold epistemic confidence and justification scores from low-resource demographic profiles. We argue that current safety alignment acts as a paternalistic filter, transforming conversational AI into agents of narrative segregation -- a manifestation of \emph{hermeneutical injustice} in Fricker's~\cite{fricker2007} sense that demands urgent pedagogical auditing.
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning
Ahmad D. Suleiman, Daqing Hou, Maliha Noushin Raida
Design problems (DPs)—scenario-based prompts requiring application of project concepts in new contexts—offer a way to assess higher-order thinking in project-based learning. Surveys of 31 instructors and evaluation of 80 LLM-generated DPs showed that instructors value DPs but find creation time-consuming, while LLMs can generate high-quality prompts with strong expert agreement. Student performance on DPs correlated weakly with traditional grades, indicating DPs capture distinct aspects of higher-order thinking and can complement conventional assessments.
Original abstract
Project-based learning (PjBL) is common in computing education, but traditional assessments of PjBL often fail to capture higher-order thinking (HOT), especially in transfer contexts. This study introduces "design problems" (DPs): concise, scenario-based prompts that require applying project concepts in new situations, to address this gap. We examined instructor perceptions, the ability of large language models (LLMs) to generate DPs, and student experiences. Surveys of 31 instructors, evaluation of 80 LLM-generated DPs, and student performance data showed that while instructors value DPs, creation effort is a barrier. LLMs helped by producing high-quality prompts with strong expert agreement. Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs may capture distinct aspects of HOT. Keystroke data also suggested deeper cognitive engagement of students through planning and revision behaviors. Overall, DPs appear to be a useful complement to traditional assessments, especially in situations where AI use or collaboration may undermine individual learning.
- PaperJournal of Second Language Writing13 Jul 2026
Understanding L2 teachers’ expertise in the design and implementation of digital multimodal composing
Zhenhao Cao, Zhicheng Mao
This study examines how L2 teachers develop and apply expertise when designing and implementing digital multimodal composing activities, focusing on their pedagogical decisions and practices.
- PaperReCALL13 Jul 2026
Fostering generative AI-supported self-regulated learning (GenAI-SRL) in informal digital language learning through literacy and interactions: A two-stage PLS-SEM-ANN approach
Xiaoqi Wang, Lawrence Jun Zhang
This study examined how GenAI literacy and interactions affect self-regulated learning (SRL) in informal digital language learning among 343 Chinese university learners. Using PLS-SEM and ANN, awareness and evaluation of GenAI significantly predicted SRL, while usage and ethics did not; student–student, student–teacher, and student–GenAI interactions all facilitated SRL, with student–student interaction being the strongest predictor.
Original abstract
Generative artificial intelligence (GenAI) enables foreign language learners to extend their learning beyond formal instruction and develop their autonomy. However, research has not adequately examined how learners regulate their learning with GenAI or how their GenAI literacy and multiple types of interactions influence their self-regulated learning (SRL) in GenAI-supported informal digital language learning settings. We address this gap by analyzing data from 343 Chinese university foreign language learners through partial least squares structural equation modeling (PLS-SEM) and artificial neural networks (ANN). PLS-SEM showed that awareness and evaluation significantly predicted GenAI-supported SRL (GenAI-SRL), whereas usage and ethics did not. Student–student, student–teacher, and student–GenAI interactions emerged as facilitators of GenAI-SRL. These three interaction types also significantly influenced most GenAI literacy dimensions, with three of them predicting awareness, usage, and evaluation, while only student–student and student–GenAI interactions significantly predicted ethics. Mediation analysis demonstrated that awareness and evaluation partially mediated the effects of student–student and student–teacher interactions on GenAI-SRL. The mediating pathways through student–GenAI interaction were not significant. ANN models identified student–student interaction as the strongest predictor of GenAI-SRL. These findings inform GenAI literacy development and the design of systems to support GenAI-SRL in informal learning contexts.
- PaperComputers & Education13 Jul 2026
Rewiring the knowledge field of artificial intelligence in education: A multi-level theoretical perspective on cross-level topic evolution, 2000–2024
Chuang Yang, Zeqing Xu, Yu Lin, Yu Zhang
The study proposes a multi-level theoretical perspective to analyze the evolution of AI in education topics from 2000 to 2024, focusing on cross-level topic interactions.
- PaperReCALL13 Jul 2026
The effectiveness of computerized dynamic assessment in improving L2 performance: A three-level meta-analysis
Qi Lu, Mengqi Chen, Lianrui Yang, Shaofeng Li et al.
A three-level meta-analysis of 35 studies found that computerized dynamic assessment (C-DA) significantly improves L2 performance, with effect sizes of g = 2.120 for cake format (mediation embedded in the test) and g = 1.676 for sandwich format (mediation between pretest and posttest). Moderator analyses revealed that the number of items, test content, and learners' first language influence C-DA effectiveness.
Original abstract
The growing body of research on the effects of computerized dynamic assessment (C-DA) on second language (L2) learning underscores the need for a comprehensive research synthesis to identify future research directions and inform the application of C-DA in L2 educational contexts. This meta-analysis employed a three-level modeling approach to examine the effectiveness of C-DA in improving L2 learners’ performance. It synthesized 27 effect sizes from cake format designs, in which mediation is embedded within the test sequence, and 24 effect sizes from sandwich format designs, where mediation is delivered between a pretest and a posttest, across 35 studies published between 2000 and May 27, 2025. This study also investigated the key variables that moderate C-DA effectiveness. Findings reveal large, significant positive effects of both the cake and sandwich formats on L2 performance improvement (cake format: g = 2.120, p < .001; sandwich format: g = 1.676, p < .001), with the cake format tending to yield larger effect sizes. This may be because the cake format captures gains during mediation, whereas the sandwich format reflects post-mediation outcomes. Moderator analyses show that the number of items, test content, and learners’ first language affect C-DA effectiveness in promoting L2 performance. Drawing on the synthesized findings, this study contributes to theoretical, methodological, and technological understandings of C-DA and offers suggestions for future research in this domain.
- PaperRELC Journal13 Jul 2026
Conceptualizing and assessing competence in English as a lingua franca context
Hyejeong Kim
This thematic review argues that current assessments fail to capture English as a lingua franca (ELF) communication and proposes using indigenous criteria from communities of practice. It advocates for a strong approach focusing on real-world task performance rather than language-specific aspects, and calls for evidence gathering on what experienced ELF users value in communication.
Original abstract
While the field of English as a lingua franca (ELF) has been growing, there is currently no assessment that fairly captures the nature of ELF communication. The necessity for such an assessment has been emphasized by some scholars, but concrete efforts to develop one have been slow. To stimulate discussion and evidence gathering, this thematic review examines ELF within the framework of communities of practice, highlighting key components of the theory to underscore the importance of learning in and through practice within these communities. The discussion then turns to indigenous criteria − that is what members of communities of practice actually value when evaluating or assessing their peers − which is highly relevant in English for specific purposes assessment. In fact, many ELF studies can be classified as contexts for English for specific purposes. By reviewing studies that explore indigenous criteria, we can gain a better understanding of how domain or non-language specialists determine what is important for performance or communication within their respective communities, as well as how these criteria differ from those perceived by language specialists or applied linguists. A strong approach (i.e., focusing on task performance in real-world contexts) is then suggested, as a weak approach (i.e., focusing on language-related aspects) is inadequate for assessing ELF interactions, based on characteristics described in both ELF studies and research on indigenous criteria in various English for specific purposes contexts. The discussion concludes with a call for evidence gathering on what experienced ELF users perceive as important for communication in ELF settings.
- PaperarXiv — AI in Education (cs.CY)12 Jul 2026
Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications
Nasser Giacaman, Valerio Terragni, Paul Denny, Viraj Kumar
Analyzed four years of undergraduate programming submissions to understand how students write natural-language comments as specifications for AI code generation. Introduced a taxonomy of comment types, code expression levels, and code constructs, finding that students predominantly wrote 'What' comments and shifted toward 'How' comments for procedural tasks, focusing more on verifying generated code than iterating on comments.
Original abstract
As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and rely on these tools to generate code, shifting emphasis from code writing to specification. Yet little is known about the comments students write as specifications in AI-assisted programming tasks. We analyze a four-year dataset of undergraduate programming submissions and reflections from tasks in which students wrote comments to guide code generation and refined solutions using test-case feedback. We introduce a taxonomy spanning three dimensions: comment type, code expression level, and code construct. Using automated classification, we examine how these dimensions vary across attempts and how students describe the process in their reflections. Our findings show that students mostly wrote natural-language What comments, shifted toward How comments for more procedural constructs, and focused more on verifying generated code than on repeatedly rewriting comments.
- PaperComputer Assisted Language Learning12 Jul 2026
Engagement without attainment? Exploring AI-assisted language learning through SLA theories
Murod Ismailov, Yuichi Ono, Thomas K.F. Chiu
Explores AI-assisted language learning through second language acquisition (SLA) theories, questioning whether high learner engagement with AI tools necessarily leads to actual attainment.
- PaperTESOL Quarterly12 Jul 2026
“Disciplinary Writing Is Like Playing with Legos”: Exploring a Multilingual International Student's Navigating Science Writing
Hongye Zeng, Laura Mahalingappa
A three-year case study of a multilingual international student reveals he used translingual, multimodal, and interdisciplinary strategies—including GenAI—to navigate science writing at a US university, shifting from viewing it as high-stakes evaluation to an agentive identity process. The findings advocate for linguistically responsive pedagogy that treats writing as a space for disciplinary engagement and identity formation, not just assessment.
Original abstract
This study explores the experiences of an undergraduate multilingual international student in becoming a science writer at a large public university. Taking the lenses of identity and investment, language socialization and multiliteracies, we address the question: How does a college‐aged multilingual science student navigate disciplinary writing over time? Interviews using talk around text based on the student's disciplinary writing samples over 3 years tracked his unique learning progress. Findings reveal he employed strategies including translingual, multimodal, and interdisciplinary practices and used GenAI to participate in disciplinary writing tasks, particularly when instruction and feedback were limited or harmful. The student gradually shifted from seeing disciplinary writing as a high‐stakes evaluation to understanding it as an agentive process of enacting future professional identity. By centering the lived experience of a multilingual writer, the study offers insights into the challenges and possibilities of disciplinary writing development in higher education (HE) classrooms. It provides a window into how students demonstrate resilience and agency to support learning, build disciplinary knowledge, and envision professional futures. The findings call for more linguistically responsive pedagogical practices in HE science, technology, engineering, and mathematics fields that go beyond writing as a tool for assessment, but rather a vital space for disciplinary engagement and identity formation.
- PaperTESOL Quarterly12 Jul 2026
Language Teacher Educator Identity, Wellbeing and Agency Nested in a Community of Practice: A Collaborative Autoethnography
Yasemin Tezgiden‐Cakcak, Aycan Demir‐Ayaz, Işıl Günseli Kaçar
Three language teacher educators in Türkiye used collaborative autoethnography to examine how their community of practice (CoP) helped them navigate identity tensions, boost wellbeing, and enhance agency. The study reconceptualizes the identity-wellbeing-agency triangle as a mutual, recursive process nested in a CoP, showing how such spaces empower educators in neoliberal academia.
Original abstract
In this collaborative autoethnography, we, as three university‐based language teacher educators in Türkiye, explore how our community of practice (CoP) acted as a catalyst in navigating our identity tensions, boosting our wellbeing and contributing to our agency. Grounded in Wenger's (2009) framework of CoP, we embarked on a critical journey of navigating the tensions between our aspired and assigned identities as teacher educators in a dialogic and reflective third space of a CoP. The collegial space of our CoP enabled us to enact our agency and nurtured our wellbeing, embracing us in a safety net of solidarity and empowering us professionally in neoliberal academia. Conceptualizing the identity–wellbeing–agency triangle as a mutual, recursive and interdependent process nested in a CoP, this study showcases how LTEs reclaim their legitimacy, nourish their wellbeing and enact their agency in a safe dialogic space. Such a reconceptualization may bring new insights to LTEs in different contexts.
- PaperComputers and Education: Artificial Intelligence11 Jul 2026
Fair and explainable educational recommendations with a hybrid Graph-GRU framework
Edmund Evangelista, Syed M. Salman Bukhari
The study introduces a Hybrid Heterogeneous Knowledge Graph-Gated Recurrent Unit framework that combines graph embeddings with sequential modeling to provide fair, robust, diverse, and explainable educational recommendations. It integrates multi-objective training, reranking, and model-centric explainability, achieving strong predictive performance on a Moodle LMS dataset while improving intra-list diversity and moderate catalogue coverage. However, popularity bias remains evident, highlighting the need for continued attention to fairness.
Original abstract
Artificial Intelligence (AI) recommender systems are increasingly used in education to personalize learning and help students navigate large collections of digital learning resources. However, many existing approaches emphasize predictive accuracy over fairness, robustness, diversity, and transparency. This creates an important educational challenge. The students with limited participation histories may receive less reliable support, while highly popular resources may dominate recommendation lists and limit access to other useful learning materials. To address this challenge, this study aims to develop and evaluate a responsible educational recommender framework that supports personalized learning resource navigation while making recommendation behaviour more fair, stable, diverse, and explainable. This study introduces the Hybrid Heterogeneous Knowledge Graph-Gated Recurrent Unit (Hybrid HKG-GRU) framework, which combines heterogeneous graph embeddings with sequential modelling to capture both the relational structure of course materials and the temporal dynamics of learner interactions. The framework integrates three contributions: (i) multi-objective training with Group Distributionally Robust Optimization (GroupDRO), (ii) Maximum Marginal Relevance (MMR) reranking to reshape exposure patterns, and (iii) built-in, model-centric explainability through path-based and counterfactual analyses. The empirical evaluation shows strong predictive performance with HR@10 = 0.68 and MRR = 0.41 on Moodle LMS logs dataset that comprises of 152 students, 59 resources, and approximately 150k interactions. The model also achieves high intra-list diversity and moderate catalogue coverage, while showing moderate counterfactual stability for many learners (median CR@10 = 1.0), although catalogue-level popularity bias remains evident. The framework provides model-centric interpretability and verification intended to support more transparent educational recommendation. The study positions the framework as a technically auditable approach for improving how learning resources are recommended, inspected, and monitored in educational settings. By integrating fairness, robustness, and explainability as coequal design objectives, the Hybrid HKG-GRU provides a methodological foundation for more responsible and accountable recommender systems in education and related high-stakes contexts.
- PaperComputers & Education10 Jul 2026
Educational prompt engineering self-efficacy scale (Ed-PESS): Development and psychometric validation
Fatih Karataş, Recep GÜR, Barış Eriçok, Fatma BAŞARIR et al.
A scale measuring educational prompt engineering self-efficacy was developed and psychometrically validated, providing a tool to assess educators' confidence in crafting effective prompts for AI systems.
- PaperTESOL Quarterly10 Jul 2026
Handwritten Versus Typed Notes: The Impact of Note‐Taking Modes in Second Language Listening Tests
Jieun Kim
The study compared handwritten, typed, and no note-taking conditions among 305 Korean L2-English adults on TOEFL listening tasks. No significant overall score differences were found, but item-level DIF analysis revealed mode effects for lower-ability participants on two items. Handwritten notes contained more information units and nonlinguistic elements, suggesting test developers should reconsider note-taking policies for validity and fairness.
Original abstract
Technological advancements have influenced note‐taking practices in classrooms as well as their treatment in standardized language assessments. Many high‐stakes English listening tests (e.g., TOEFL iBT, IELTS) include varying note‐taking guidelines, often established without empirical support. This study examined the effects of note‐taking modes on listening performance and note content. A total of 305 L1‐Korean L2‐English adults were randomly assigned to handwriting ( n = 102), typing ( n = 102), or no note‐taking ( n = 101) conditions and completed TOEFL iBT lecture comprehension tasks. Linear regression results revealed no significant differences in overall test scores across modes. Uniform Differential Item Functioning (DIF) analyses indicated that one item favored handwriting and another favored typing, while non‐uniform DIF emerged for the same items only among lower‐ability participants. When examining note‐taking features, no significant differences were found in word count or translanguaging across modes. However, handwritten notes contained more information units, verbatim transcription, and nonlinguistic elements. The overall lack of significant performance differences, together with finer‐grained differences at the item and note‐content levels, calls for test developers to reconsider note‐taking policies in terms of test validity and fairness. Pedagogical implications are discussed, highlighting the importance of granting students autonomy to choose note‐taking modes and providing instructional sessions to experience various note‐taking strategies.
- PaperELT Journal10 Jul 2026
Teacher code alternation in a Sri Lankan tertiary level EMI classroom
Dilini Leelachandra, Dushyanthi Mendis
Teacher code alternation in an English medium instruction (EMI) classroom at a Sri Lankan university serves five distinct functions: enhancing comprehension, managing classroom dynamics, building rapport, alleviating English-related stress, and clarifying complex ideas. The study underscores code alternation as a strategic linguistic tool for inclusive multilingual education.
Original abstract
This study explores teacher code alternation in an English medium instruction classroom at a higher education institution in Sri Lanka. Acknowledging that code alternation is common in multilingual settings, this case study highlights its various functions, motivations, and implications for classroom purposes, addressing a gap in research within the Sri Lankan higher education context. Data were collected through classroom observations, transcriptions of recorded lessons, and semi-structured interviews. The analysis revealed five distinct functions of code alternation that aid the learning process. As a strategic linguistic tool, code alternation enhances comprehension, manages classroom dynamics, and builds rapport. Specifically, it serves to alleviate the stress of using English, foster a comfortable environment, and clarify complex ideas, thereby underscoring the importance of teacher–student perspectives on the use of code alternation for inclusive, multilingual education.