Topics
- PaperAssessing Writing18 Jul 2026
Modeling the reading-to-writing pipeline: Knowledge graph and LLM-based assessment framework for source-based writing
Byungyeon Yun, Miranda Moe, Lauren E. Flynn, Püren Öncel et al.
This study proposes a framework that models the reading-to-writing pipeline using knowledge graphs and large language models to assess source-based writing. The approach aims to capture how writers integrate source information and improve automated evaluation of writing quality.
- PaperBritish Journal of Educational Technology18 Jul 2026
Not a universal benefit: Examining the differential effects of emotional AI on L2 pre‐service teachers' language learning
Zhuo Wang, Hui Pang
A quasi-experimental study with 147 pre-service teachers found no overall benefit of emotional AI for L2 vocabulary learning, and lower-proficiency learners actually showed better attitudes and intrinsic motivation with a regular AI agent. Qualitative analysis identified four learner archetypes (e.g., Constructive Inquirer, Attentive Pupil) that moderate the effect, suggesting that verbose emotional scaffolding can increase cognitive load under high task difficulty.
Original abstract
While emotional design in educational AI is often presented as a universal benefit, this study challenges that assumption, investigating when, for whom and how it impacts L2 vocabulary learning. This paper reports on the first phase of a larger project, analysing data from a quasi‐experimental study with 147 pre‐service teachers who interacted with either an emotional or a regular AI agent. Data included pre‐/post‐tests, questionnaires and complete AI chat histories. Quantitative analysis revealed no overall difference in vocabulary acquisition or other affective variables. A more nuanced pattern emerged in an exploratory subgroup analysis: contrary to prevailing assumptions, the regular agent—not the emotional one—preserved significantly better learning attitudes ( p = 0.001) and intrinsic motivation ( p = 0.038) for learners with lower baseline proficiency. A possible mechanism is suggested by a significant negative correlation between self‐reported cognitive load and vocabulary scores, found only in the emotional AI group: under high task difficulty, the emotional agent's verbose scaffolding may have added extraneous cognitive load rather than serving as a supportive buffer. We treat this correlational evidence as tentative rather than demonstrated. To explain these divergent outcomes, a qualitative analysis of interaction patterns identified four distinct learner archetypes—from the high‐agency ‘Constructive Inquirer’ to the passive ‘Attentive Pupil’. The findings indicate that the emotional AI's efficacy might be moderated by these profiles; its verbose emotional scaffolding may have added extraneous processing demands for the vulnerable ‘Attentive Pupil’ while being perceived as an inefficient frustration by the task‐oriented ‘Demanding Critic’. The study concludes that a one‐size‐fits‐all approach to emotional AI is suboptimal. Practitioner notes What is already known about this topic Emotional design in digital learning environments is widely considered beneficial for enhancing learner engagement, motivation and positive emotions. Affective AI systems are increasingly being developed for education with the goal of providing personalized and psychologically supportive instruction. Individual learner differences, such as prior knowledge, motivation and confidence, are known to be significant factors that impact the effectiveness of any educational intervention. What this paper adds This paper demonstrates that the benefits of emotional AI are not universal. Its efficacy is highly conditional, with the regular agent—not the emotional one—preserving better learning attitudes and motivation for lower‐proficiency learners under high cognitive load conditions. It tentatively provides a new explanatory framework of four learner archetypes (the Constructive Inquirer, Demanding Critic, Attentive Pupil and Disengaged Bystander) based on observable interaction patterns (agency and questioning effectiveness), which explains why different learners react divergently to the same AI design. It identifies a cognitive load threshold boundary condition: under high task difficulty, verbose emotional scaffolding may exceed learners' working' memory capacity and act as an extraneous processing burden rather than affective support, consistent with experimental evidence on the moderating role of task difficulty in emotional design. It reveals a potential design tension: emotional support is not uniformly beneficial—verbose affective scaffolding intended to help a passive learner (the ‘Attentive Pupil’) may instead add extraneous cognitive load under demanding tasks, while being perceived as an inefficient frustration by a high‐agency, task‐oriented learner (the ‘Demanding Critic’). Implications for practice and/or policy The design and implementation of educational AI must move beyond a ‘one‐size‐fits‐all’ model. Practitioners should select and advocate for tools that can adapt their emotional persona based on user needs, rather than applying
- PaperLanguage Teaching Research18 Jul 2026
The Effect of Text Shadowing on English Language Learners’ Pronunciation Development: A Quasi-Experimental Study
Mishelle Kehoe-Seamons, Mark Tanner, K. James Hartshorn, Rob Martinsen
A ten-week quasi-experimental study with intermediate adult ELLs found that both text shadowing and control groups significantly improved in fluency, comprehensibility, and accentedness, but no significant difference emerged between the groups on any measure. Participants in the shadowing group provided positive qualitative feedback, suggesting perceived value in oral communication curricula.
Original abstract
Shadowing is a technique that has been shown to significantly improve English language learners’ (ELLs) oral fluency and comprehensibility. However, previous research showing dramatic gains over time for ELLs has varied in design, with some studies including control groups and others not, thus complicating interpretations of efficacy. These studies have also been conducted in English and other languages with attention to advanced-level ELLs and beginning-level ELLs. Additional research is needed to provide insight into the effect of shadowing on intermediate learners, as well as using treatment and control groups to get a more accurate picture of identified changes in fluency, comprehensibility, accentedness, and imitative ability. A ten-week quasi-experimental study of text shadowing practice was conducted with intermediate-level adult ELLs studying in a large western university’s intensive English program (IEP). Speech samples from the pretest and posttest were rated by non-expert native English-speaking raters for fluency, comprehensibility, accentedness, and the quality of imitative speech on nine-point Likert scales, as has been used in other pronunciation studies. A repeated-measures ANOVA showed that all participants improved significantly from the pretest to posttest in fluency and comprehensibility, with a reduction in their accentedness. While raw gain scores tended to be higher in the treatment group than the control group, the analysis showed no statistical difference between the groups in any of the four measures. The treatment group participants did share positive qualitative feedback regarding shadowing, emphasizing the value of this dynamic activity in an oral communication curriculum.
- NewsLanguage Magazine17 Jul 2026
Bilingualism Delays Brain Aging By 6 Years
New research presented at the FENS Forum 2026 indicates that bilingualism can delay brain aging by up to six years. The study, led by Dr Lucia Amoruso, found that speaking multiple languages, especially when acquired early, is associated with younger brain age.
Original abstract
New research suggests that the more languages people speak, the younger their brains appear. Learning more languages, especially when younger, and achieving fluency in a second language also seem to slow brain aging. The research was presented at the recent Federation of European Neuroscience Societies (FENS) Forum 2026 by Dr Lucia Amoruso from the Basque […] The post Bilingualism Delays Brain Aging By 6 Years appeared first on Language Magazine .
- NewsBritish Council TeachingEnglish17 Jul 2026
How to use AI to support reflection and autonomy in teacher education – webinars
This news announces webinars on using AI to foster reflection and autonomy in teacher education, hosted by TeachingEnglish.
Original abstract
<span class="field field--name-title field--type-string field--label-hidden">How to use AI to support reflection and autonomy in teacher education – webinars</span> <span class="field field--name-uid field--type-entity-reference field--label-hidden"><article class="user-profile-card profile"> <div class="user-profile-card-inner"> <div class="avatar"> </div> <div class="username-row"> <div class="username h6"> TeachingEnglish </div> </div> <div class="profile-bio small"> <p>This resource was developed by the TeachingEnglish editorial team.</p> </div> </div> </article> </span> <span class="field field--name-created field--type-created field--label-hidden"><time datetime="2026-07-17T19:28:11+00:00" title="Friday, July 17, 2026 - 19:28" class="datetime">Fri, 07/17/2026 - 19:28</time> </span> <div class="layout layout--onecol"> <div class="layout__region layout__region--content"> <div class="magazine-block-hero-content block block-layout-builder block-field-blocknodemagazinefield-hero-content"> <div class="break-out hero-bg"> <div class="container-fluid px-0"> <div class="field field--name-field-hero-content field--type-entity-reference-revisions field--label-hidden field__items"> <div class="field__item"> <div class="paragraph paragraph--type--image paragraph--view-mode--default"> <img loading="lazy" src="https://www.teachingenglish.org.uk/sites/teacheng/files/styles/wide_1920_x_700_focal_point_crop/public/Getty%201781988444.jpg?h=9eb0d413&itok=aixuaJPZ" width="1920" height="700" alt="A person uses a smartphone while standing next to a large blue-lit digital display. The scene represents engagement with digital technology and online information." class="img-fluid image-style-wide-1920-x-700-focal-point-crop"> </div> </div> </div> </div> </div> </div> <div class="hidden block block-system block-system-breadcrumb-block"> <nav role="navigation" aria-labelledby="system-breadcrumb"> <h2 id="system-breadcrumb" class="visually-hidden">Breadcrumb</h2> <ol class="breadcrumb"
- PaperEdArXiv (OSF Preprints)17 Jul 2026
Designing and Evaluating an Integrated AI-Based Educational System for Enhancing Critical Analysis Skills in Pre-Service Teachers
Hossein Talebzadeh
An integrated AI-based educational system was designed and evaluated to enhance pre-service teachers' critical analysis skills. A qualitative case study with history and social science teachers revealed dual-layered technical and human challenges, but also showed that AI can act as a 'pedagogical assistant' by engineering cognitive conflict to foster transformative learning. The effectiveness of AI depends on its integration into a structured educational system that reframes challenges as learning opportunities.
Original abstract
Given the emerging challenges and opportunities of generative artificial intelligence (GenAI) in teacher education, this study designs, implements, and evaluates an integrated educational system aimed at enhancing the critical analysis skills of pre-service teachers. This qualitative case study was conducted with the participation of history and social science pre-service teachers at Farhangian University. Data documenting their experiences within an AI-assisted content analysis project was collected via team and individual evaluation forms and analyzed using the thematic analysis method. The findings revealed dual-layered technical and human challenges, alongside multifaceted technical, pedagogical, and soft skill learning outcomes. More importantly, the results demonstrated a dialectical relationship between challenge and learning, conceptualizing the role of AI as a "pedagogical assistant" that fosters transformative learning by deliberately engineering cognitive conflict. We conclude that the effectiveness of AI in teacher education depends on its integration into a meticulously structured educational system—one where challenges are systematically reframed as learning opportunities and technology serves as a catalyst for shaping the professional identity of innovative educators.
- PaperEdArXiv (OSF Preprints)17 Jul 2026
Distributed Auditing of Cognitive Balance in History Textbooks: Application of the CAP Protocol within a Pre-Service Teacher Network
Hossein Talebzadeh
Analyzed the cognitive balance of Iranian upper secondary history textbooks using the Comparative Adjudication Protocol (CAP) with 43 pre-service teachers. Found extremely low student engagement indices (text involvement 0.019, image involvement 0), with 96.7% of lessons classified as passive 'banking education'. The CAP protocol achieved 86.4% inter-rater agreement, demonstrating reliability for content analysis.
Original abstract
Background: Despite theorists like Sam Wineburg emphasizing the development of historical thinking, history textbooks in centralized educational systems continue to be dominated by rote-learning approaches, leaving the systematic evaluation of their cognitive balance—particularly at the upper secondary level—largely neglected. Objective: This study aims to analyze the cognitive balance of upper secondary history textbooks (History 1, 2, 3, and Contemporary History of Iran) and evaluate the efficacy of the Comparative Adjudication Protocol (CAP) within a network of human auditors. Methodology: Engaging 43 pre-service teachers divided into 15 independent research teams, this study analyzed the content of 4 textbooks comprising 61 independent lessons using the William Romey technique and the CAP protocol. The process was executed across four distinct phases: capacity building, independent AI-assisted auditing, blind peer auditing, and discrepancy adjudication. Findings: The overall text involvement index was exceptionally low (0.019), far below the active learning threshold of 0.4, with 96.7% of the lessons falling into the passive "banking education" category. The image involvement index was absolute zero (0), and while the question involvement index was 1.29 (with an unbalanced distribution ranging from ∞ to 0.06), it revealed a severe structural gap between passive textual content and active analytical expectations. The CAP protocol demonstrated high reliability, achieving an 86.4% final inter-rater agreement and yielding a 12.5% increase in initial consensus. The dominant discrepancy pattern (A vs. B at 48.3%) indicated that distinguishing "objective historical facts" from "author interpretation" remains the primary analytical challenge for human coders in historical texts. Conclusion: Despite their narrative essence, high school history textbooks exhibit a highly passive cognitive structure lacking a coherent strategy to foster active student engagement. Implementing the CAP protocol within the R2A-TACI-PACT conceptual framework underscores the vital necessity of a "Human-in-the-Loop" (HITL) approach in historical content analysis. Curriculum revision and empowering pre-service teachers with algorithmic auditing literacy are highly recommended.
- PaperEdArXiv (OSF Preprints)17 Jul 2026
The Biological Meaning of Surface Area-to-Volume Ratio: A Conceptual Framework for Modeling and Teaching bio-SA/V
Tzachi Bar
Introduces the concept of bio-SA/V as a biological trait scaling with surface area-to-volume ratio, and presents a model distinguishing obligatory and pseudo bio-SA/V functions. The model reveals cognitive challenges for learners and common explanatory errors in biology education, offering new teaching and assessment approaches.
Original abstract
Surface area–to–volume (SA/V) ratio is widely used to explain scaling of biological phenomena. Despite its central role in biology and science education, the biological meaning of SA/V has remained undefined. This paper introduces the concept of bio-SA/V, defined as a biological trait whose magnitude scales with the SA/V of the body producing it, and presents a model of bio-SA/V components and their relationships. The model identifies obligatory bio-SA/V, in which a biological function emerges from two components: one occurring at the body surface and proportional to surface area, and the other occurring within the body volume and proportional to the body's volume. Also defined and characterized is pseudo bio-SA/V, describing traits that scale with the characteristic length of the body, the inverse of SA/V. The value of the new concept is demonstrated through simplification of the explanation of erythrocyte deformability. Analysis of the model reveals that understanding obligatory bio-SA/V functions requires coordinating multiple layers of proportional reasoning and synthesizing surface- and volume-based components into a composed biological function, posing a significant cognitive challenge for learners. The model also reveals common explanatory problems in the literature, including oversimplified explanations that omit the volume component and overly complicated explanations that could be reduced to characteristic length. These insights have implications for biology education. The framework suggests new approaches for teaching and assessing SA/V-based reasoning and clarifies, for the first time, the appropriate use of units in SA/V-based explanations.
- PaperEdArXiv (OSF Preprints)17 Jul 2026
Why AI Implementation Fails in Schools: The Human Readiness Gap
Dr NIRMALA KRISHNAN
Existing theoretical models for AI implementation in schools fail to explain recurring failures because they were designed for fixed, bounded change episodes. Using case studies of two large-scale LAUSD initiatives, this paper identifies a 'Human Readiness Gap' and a causal 'Human Failure Chain' that current frameworks overlook. It argues that organizational and relational conditions, not just funding and technology, determine adoption success.
Original abstract
This paper conducts a critical audit of the theoretical models most often used to plan, justify, and explain artificial intelligence (AI) implementation in schools, and asks why none of them adequately explains the pattern of failure that recurs across technology generations. Individual-level acceptance theories (the Technology Acceptance Model; the Unified Theory of Acceptance and Use of Technology), diffusion theory, organization-level adoption frameworks (the Technology-Organization-Environment framework), sequential change-management models (Kotter's eight-step process), and implementation science (the Consolidated Framework for Implementation Research; the National Implementation Research Network's implementation-driver model) each illuminate part of the adoption process but were built for change episodes with properties AI implementation in schools does not share: a fixed protocol to be delivered with fidelity, a bounded start and end point, or a rational individual decision-maker as the unit of analysis. Using two documented, procurement-sound, well-resourced AI and technology initiatives at the same U.S. school district that nonetheless collapsed the Los Angeles Unified School District's 2013 one-to-one iPad programme and its 2024 “Ed” AI chatbot this paper traces a recurring causal sequence, termed here the Human Failure Chain, that existing models do not name. It argues that a Human Readiness Gap exists between what districts typically assess before AI adoption (funding, devices, vendor credentials, policy compliance) and the organizational and relational conditions that determine whether adoption succeeds. The paper does not claim to have tested this causal sequence statistically; it offers a critical synthesis and an illustrative case analysis and proposes the comparative and process-tracing studies that would be required to establish the Human Failure Chain as a validated causal model. Keywords: AI implementation failure; educational technology; implementation science; organizational readiness; technology adoption theory; school transformation
- PaperEdArXiv (OSF Preprints)17 Jul 2026
Introducing Human Readiness for AI in Education: A New Theoretical Framework for School Transformation
Dr NIRMALA KRISHNAN
The paper introduces a new theoretical framework, Human Readiness for AI in Education, arguing that the success of AI implementation in schools depends primarily on the readiness of people rather than the technology itself. It synthesizes organizational readiness theory, technology-acceptance research, and sociotechnical systems to propose five dimensions of human readiness: shared purpose, professional capability, relational trust, collective judgement, and adaptive growth. The framework is positioned against established adoption theories and includes a research agenda for empirical validation.
Original abstract
Three decades of large-scale technology investment in schools from one-to-one laptop programmes to interactive whiteboards to, most recently, artificial intelligence (AI) has produced a consistent pattern: substantial capital expenditure paired with modest, uneven, or absent instructional change. Prevailing accounts of this pattern locate the explanation in the technology itself its cost, its usability, its alignment with curriculum and prescribe better tools, more training, or clearer policy as the remedy. This paper argues that the explanation lies elsewhere. Synthesizing organizational readiness theory, technology-acceptance research, and the sociotechnical-systems tradition with an emerging body of evidence on AI adoption in K–12 settings, the paper introduces Human Readiness for AI in Education as a conceptual framework in which the decisive variable in AI-implementation outcomes is not the technology adopted but the readiness of the people asked to adopt it. Human readiness is proposed as a multidimensional organizational capacity spanning shared purpose, professional capability, relational trust, collective judgement, and adaptive growth that determines whether a given technology amplifies or degrades educational practice. The paper positions this construct against established adoption theories, states its core propositions and boundary conditions, and specifies the evidence that would validate or falsify it. It closes with a research agenda through which the construct can be tested. Consistent with the norms of theory-building scholarship, this paper offers a conceptual foundation, not a validated instrument or an implementation methodology. Keywords: human readiness; artificial intelligence in education; technology adoption; educational leadership; organisational readiness for change; school transformation
- PaperEdArXiv (OSF Preprints)17 Jul 2026
Human Readiness as the Missing Variable in AI Adoption in Education
Dr NIRMALA KRISHNAN
This paper positions the construct of Human Readiness for AI in Education against seven established technology adoption models, arguing that it offers a new organization-level perspective spanning dimensions not covered by existing accounts. It outlines discriminant-validity tests to empirically distinguish Human Readiness from related constructs. The paper contributes to theory development for AI implementation in schools.
Original abstract
A theoretical contribution to organizational scholarship requires more than the introduction of a new label for a familiar phenomenon; it requires showing what factors an existing account omits, how those factors relate to one another, and why the resulting explanation improves on what came before (Whetten, 1989). This paper undertakes that task for Human Readiness for AI in Education, a construct introduced in a companion paper (Krishnan, 2026a) as the primary determinant of whether artificial intelligence (AI) implementation benefits a school. Rather than introducing new theory, this paper positions Human Readiness systematically against seven established accounts of technology adoption and organizational change the Technology Acceptance Model, the Unified Theory of Acceptance and Use of Technology, Diffusion of Innovations theory, Technological Pedagogical Content Knowledge, the Substitution-Augmentation-Modification-Redefinition model, the Concerns-Based Adoption Model, and Weiner's theory of organizational readiness for change comparing each on level of analysis, outcome variable, core constructs, and explanatory scope. It argues that Human Readiness is not a relabeling of any single existing construct, but a candidate organization-level construct whose five proposed dimensions span territory these accounts divide across separate literatures. Because the value of that claim rests entirely on demonstrating that Human Readiness is empirically distinct from its nearest neighbors, the paper closes by specifying, using established construct-validation methodology (Cronbach & Meehl, 1955; Campbell & Fiske, 1959; Fornell & Larcker, 1981; MacKenzie, Podsakoff, & Podsakoff, 2011), the discriminant-validity tests the construct must pass and the outcome under which it should be judged redundant and retired. Keywords: human readiness; construct validity; technology acceptance; organizational readiness for change; discriminant validity; theory development
- PaperEdArXiv (OSF Preprints)17 Jul 2026
The Role of Artificial Intelligence in Green Education: Optimizing Teacher Workflow and Enhancing Pedagogical Design under Sustainable Development Pedagogy (SDP) Constraints
Hossein Talebzadeh
The study examined how integrating a sustainable development constraint into AI-assisted instructional design affects the quality of pre-service teachers' lesson plans. Using a quasi-experimental design with 28 teacher teams, the results showed significant improvement in design quality after the intervention. The findings suggest that AI-supported sustainable pedagogy transforms teachers into more strategic and reflective educational managers.
Original abstract
The present study investigates the effectiveness of implementing an educational constraint centered on sustainable development on the quality of pre-service teachers' instructional design within the framework of the Integrated AI Triad (IAT) model. A quasi-experimental approach was adopted, comparing the lesson plans of 28 pre-service teacher teams across a Baseline phase and a Sustainable Development Pedagogy (SDP) phase. The results of a paired $t$-test revealed a statistically significant improvement in the overall design quality after the intervention (t(27) = 13.78, p < 0.001), accompanied by an exceptionally large effect size (Cohen's d = 2.80). The highest surge was observed in the SDP workflow compliance index, which exhibited an increase of 1.83 units. These findings demonstrate that enforcing a rigorous, AI-assisted zero-paper resource management constraint serves as a strategic catalyst, transforming the teacher's role from a conventional pedagogical designer into an efficient, highly reflective, and strategic educational manager.
- PaperEdArXiv (OSF Preprints)17 Jul 2026
The Effectiveness of an Intensive AI Professional Development Program on Primary School Teachers' AI-PCK: The Pivotal Role of Assessment Rubric Design
Hossein Talebzadeh
An intensive AI professional development program for primary school teachers significantly improved their AI-based Pedagogical Content Knowledge, with the largest gains in assessment rubric design. The study highlights that mastering assessment rubric design is now a higher priority than content creation tools for teachers in the generative AI era.
Original abstract
This sub-study evaluates the effectiveness of an intensive Generative Artificial Intelligence (AI) professional development program in enhancing primary school teachers' AI-based Pedagogical Content Knowledge (AI-PCK) and identifies the specific component that yielded the highest level of improvement. Using a quasi-experimental pretest-posttest design, the study examined a sample of 142 in-service primary school teachers. Data were collected via a researcher-developed questionnaire covering 5 distinct components of AI-PCK and analyzed using paired t-tests and Cohen’s d effect size. The findings confirmed a highly significant overall positive impact of the professional development program on the teachers' total AI-PCK scores (p<0.001), accompanied by an exceptionally large effect size (d=3.71). Learning gain analysis revealed that the fifth component—familiarity with and application of assessment rubrics—experienced the most substantial improvement, achieving the highest learning gain (2.65) and the largest individual effect size (d=4.82). This result implies that in the era of Generative AI, the critical professional development priority for primary school teachers has shifted from content creation tools toward mastering the design of authentic, process-oriented assessment rubrics. This research offers a foundational framework for educational policymakers addressing the challenge of AI authenticity in primary education.
- PaperEdArXiv (OSF Preprints)17 Jul 2026
Evaluating the Effectiveness of Generative Artificial Intelligence in Empowering Teachers for Constructivist Instructional Design: A Case Study of the SAHAB Model
Hossein Talebzadeh
The study evaluates the SAHAB model, which integrates generative AI, social constructivism, and Iran's educational reform domains. A 12-hour intervention with 33 teachers significantly improved their instructional design competencies, with a large effect size. Qualitative findings highlight increased professional agency and cognitive augmentation, suggesting AI as a collaborative scaffold rather than a replacement.
Original abstract
The gap between the abstract goals of macro-level educational reform policy documents and classroom realities has long stood as a chronic challenge in educational transformation. This study examines the effectiveness of the "SAHAB" pedagogical model (Smart Indigenous Pedagogical System) on teachers' professional competencies in designing learning units that integrate generative AI, social constructivism, and the six educational domains of Iran’s Fundamental Reform Document (FRD). A quasi-experimental, single-group pretest-posttest design was employed. The target population comprised all teachers at the Noor-e-Iman Educational Complex during the 2025–2026 academic year. Using a census approach, all 33 eligible teachers participated in a 12-hour intervention workshop conducted over three days. Data were collected via a researcher-developed SAHAB Instructional Design Competence Questionnaire, which demonstrated high validity and reliability. The quantitative data were analyzed using paired t-tests and Cohen’s d effect size calculation. The results indicated that the mean score of teacher competencies significantly increased from 3.05 in the pretest to 4.33 in the posttest (p < 0.001). Furthermore, a robust Cohen's d effect size of 1.18 confirmed the substantial impact of the intervention. Qualitative thematic analysis of the teachers' open-ended reflections revealed three core themes: reclamation of professional agency, cognitive augmentation, and the operationalization of abstract educational standards. These findings suggest that generative AI, functioning as a collaborative cognitive scaffold rather than a replacement, can substantially reduce teachers' cognitive load, shifting their role from passive content deliverers to active designers of deep learning pathways.
- PaperarXiv — AI in Education (cs.CY)17 Jul 2026
EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha, Jungpil Shin
EduGuard, a retrieval-augmented generation (RAG) tutoring framework for introductory programming, incorporates query understanding, instructor-approved retrieval, rubric-aware generation, claim verification, and overreliance control. Compared to baselines including GPT-4o-mini, EduGuard achieves higher correctness, grounding, and rubric alignment while reducing hallucination and direct-answer leakage. A pilot study showed improved learning outcomes and reduced overreliance, suggesting that safe GenAI tutoring requires explicit pedagogical control and evidence verification.
Original abstract
Generative AI (GenAI) is increasingly used by students for programming explanation, debugging, and assignment support. Yet unrestricted large language model (LLM) tutors can hallucinate, contradict course policy, reveal complete solutions, and foster passive dependence. This paper presents EduGuard, a safe retrieval-augmented generation (RAG) tutoring framework for introductory programming. EduGuard integrates query understanding, instructor-approved course retrieval, pedagogical strategy selection, rubric-aware generation, claim-level verification, and overreliance control. To make evaluation provenance explicit, we construct BILearn-CS, a 600-query instructor-authored, TA-validated benchmark spanning concept questions, debugging cases, misconceptions, assignment-support requests, code-mixed Bangla-English queries, and adversarial direct-answer prompts. Moving beyond a synthetic-only benchmark, we further evaluate on a 150-query public CS50-style course-forum set and run a small controlled pilot with 10 undergraduates using a counterbalanced pre-test/post-test design. Using Meta-Llama-3.1-8B-Instruct as the primary generator, hybrid FAISS/BM25 retrieval, and DeBERTa-v3-large-MNLI as an architecturally separate verifier, EduGuard is compared against strong baselines: GPT-4o-mini Tutor, Llama Socratic Tutor, LPITutor-style RAG, RAG with rubric prompting, and RAG with same-model self-checking. On BILearn-CS, EduGuard attains the best correctness (90.1%), grounding (89.4%), and rubric alignment (90.8%), with the lowest hallucination (4.9%) and direct-answer leakage (9.8%). In the pilot, it raises immediate post-test accuracy from 68.4% to 81.2% and cuts overreliance from 38.0% to 17.0% relative to GPT-4o-mini Tutor. These results suggest safe GenAI tutoring requires not only retrieval or strong prompting, but explicit pedagogical control, evidence verification, and deployment safeguards.
- PaperAssessment & Evaluation in Higher Education17 Jul 2026
Navigating the moral panic: encouraging appropriate use of GenAI in the classroom rather than condemning innovation as disruption
Jennifer M. Krebsbach, Victoria L. Cross
Faculty are divided on generative AI, with some fearing it will increase cheating and reduce learning. A study of a data visualization course compared student performance across three conditions: pre-GenAI, GenAI-available, and GenAI-integrated. Results showed that when GenAI was integrated into instruction, applied performance and variability returned to pre-GenAI levels, suggesting that teaching proper use of GenAI can maintain academic integrity and authenticity.
Original abstract
Faculty are divided on how to respond to generative artificial intelligence (GenAI). Many respond in similar ways to how faculty responded to other disruptive technologies (e.g. calculators, word processors), fearing that students will cheat more and learn less. Rather than fear GenAI, we view it as the next disruptive technology that will become normalised in the classroom and examine how intentionally integrating it into instruction increases both academic integrity and authenticity. Using data from eight iterations of a lower-division Data Visualisation course, we compare student performance on two quiz types (knowledge and application) in three conditions (pre-GenAI, GenAI-available, and GenAI-integrated). In the GenAI-integrated condition, GenAI use was taught and encouraged. During the GenAI-available quarters, performance on the applied questions decreased and variability increased, indicating that only some students were using GenAI, and ineffectively. When GenAI was embedded into the assessments, applied performance and variability returned to pre-GenAI levels. Knowledge questions fell below baseline with the new format, but so did variability, potentially indicating greater equity across the quiz. By teaching students to use GenAI, we aim to support long-term learning, improve accountability, ensure authenticity, and support instructors in addressing concerns about potentially problematic technology in the classroom.
- PaperLanguage Teaching Research17 Jul 2026
Not the Artificial Intelligence, but the Sequence: Instructional Design and Oral Communication Strategy Development
Zola Chi-Chin Lai
This study compared three instructional designs in an 18-week EFL speaking course: a BOPPPS-based sequence where students used ChatGPT within a structured pedagogical frame, a flexible AI condition with self-directed ChatGPT use, and a traditional teacher-guided class without AI. The BOPPPS-based class showed the strongest gains in oral communication strategies, particularly in negotiation for meaning, fluency regulation, and accuracy monitoring, while the AI-flex class showed the smallest improvements. Results indicate that instructional design, not AI alone, drives strategic development, with AI contributing meaningfully only when embedded in a structured pedagogical sequence.
Original abstract
This study examines how instructional design shapes the development of oral communication strategies in an 18-week English-as-a-foreign-language speaking course, at a time when generative artificial intelligence is often assumed to enhance speaking performance. All three classes worked with the same curriculum and speaking tasks. One class learned through a BOPPPS (Bridge-In, Objective, Pre-Assessment, Participatory Learning, Post-Assessment, and Summary) based sequence in which students prepared ideas, participated in guided speaking activities, and reviewed their performance, with each cycle including a stage where they generated their own content before using ChatGPT to refine it. A second class received regular teacher guidance and completed the same tasks, but their use of ChatGPT was self-directed and not part of a fixed sequence (the artificial intelligence-flex condition). A third class covered the same material through teacher-supported face-to-face interaction without artificial intelligence (the traditional condition). Learners completed the Oral Communication Strategy Inventory at the beginning and end of the semester and provided reflections on how their course format influenced their speaking. The findings show a clear hierarchy in strategic development. The BOPPPS based class demonstrated the strongest gains, especially in negotiation for meaning, fluency regulation, and accuracy monitoring. The traditional class showed moderate improvements, particularly in nonverbal and social-affective strategies. The artificial intelligence-flex class displayed the smallest increases, a pattern reflected in learners’ descriptions of artificial intelligence-driven exchanges involving fewer occasions to plan, clarify, or monitor meaning. Overall, the results indicate that strategic growth is driven primarily by instructional design, with artificial intelligence contributing meaningfully only when situated within a structured pedagogical frame.
- PaperLanguage Teaching Research17 Jul 2026
When Does it Really Matter? Exploring the Conditions under which Gestures Affect Students’ Language Recall, Pronunciation Accuracy, and Cognitive Load in the EFL classroom
Ying Wang
Three controlled experiments examined how gesture type, engagement mode, and timing impact EFL learners' language recall, pronunciation accuracy, and cognitive load. Iconic gestures that were actively imitated and temporally aligned with speech improved outcomes, while misaligned gestures increased cognitive load and reduced performance. Differences between human tutors and virtual agents were modest, highlighting the need for deliberate integration of meaningful and well-timed gestures in both human and AI-mediated instruction.
Original abstract
Gestures are widely recognized as an integral aspect of multimodal communication and language learning. However, the conditions under which they effectively support learning remain unclear. This study investigated when gesture matters in English as a foreign language (EFL) learning through three controlled experiments. Experiment 1 (3 × 2 design) examined the effects of gesture type (iconic, deictic, metaphoric) and production mode (human tutor vs. virtual agent). Experiment 2 compared gesture engagement (watch vs. imitate), while Experiment 3 tested gesture timing (synchronous, gesture-first, misaligned). University-level EFL learners participated in 30-minute instructional sessions, with outcomes measured in language recall, pronunciation accuracy, and cognitive load using multivariate analysis of covariance. Results showed that gesture enhances learning selectively rather than uniformly. Higher recall and pronunciation accuracy were observed when gestures were iconic, actively imitated, and temporally aligned with speech. Misaligned gestures increased cognitive load and reduced performance. Differences between human tutors and virtual agents were present but modest. These findings extend embodied cognition by demonstrating that gesture effectiveness depends on the alignment of representational clarity, learner engagement, and gesture-speech timing. The study highlights the need for deliberate integration of meaningful and well-timed gestures in both human and language instruction mediated by artificial intelligence (AI).
- PaperComputers and Education: Artificial Intelligence17 Jul 2026
Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
Kamila Misiejuk, Sonsoles López-Pernas, Eduardo A. Oliveira, Brendan Eagan et al.
The study compares human and LLM ordered coding of qualitative text data from learners, finding systematic differences in structural, transitional, and code-level metrics. Two evaluation approaches are proposed to assess coding quality, emphasizing that LLM errors could cascade and distort automated feedback.
Original abstract
Automating the process of qualitatively coding text data from learners has been a long-standing ambition of learning analytics researchers since it represents an essential step toward delivering timely and scalable feedback. Automating this process is especially challenging in the case of ordered coding schemes —necessary for temporal analytical methods— where one text utterance can be assigned more than one qualitative code and the assignment order matters. This problem goes beyond multi-class and multi-label classification and, therefore, cannot be easily tackled using classic language models such as BERT. Recent advances in generative artificial intelligence, especially with the advent of large language models, have —allegedly— created a substantial step forward in making the goal of automatically coding complex temporal data attainable. However, little is yet known about how to implement this process in a way that most closely resembles human coding, i.e., taking into account the context in which the textual data appears for accurate interpretation. Moreover, due to the complexity of the data and its shape, the accuracy of the results cannot be computed using classic accuracy metrics. This study makes two main contributions: first, it presents two evaluation approaches for assessing the quality of ordered data coding and the usability of LLM in automatically coding ordered processes; and second, it demonstrates a method of LLM prompting that leverages a consistent context window. Our results reveal systematic and statistically significant differences between LLM and human coding across structural, transitional, and code-level metrics for binary and ordered tasks. As classification errors can propagate through automated feedback systems, relying on LLM outputs risks amplifying inaccuracies and producing misleading interpretations of learning processes.
- PaperComputers and Education: Artificial Intelligence17 Jul 2026
AI literacy-related domains and AI-TPACK readiness among preservice mathematics teachers: A factor-informed structural equation modelling study
Moeketsi Mosia, Fadip Audu Nannim, Felix Egara
A factor-informed structural equation modeling study of 130 preservice mathematics teachers in South Africa found that AI-TPACK readiness is largely unidimensional and predicted by prior AI use, critical-ethical appraisal, and support/enablers, accounting for 53% of variance. The support/enablers path had the largest coefficient but was provisional due to construct overlap. Year level was significant in the primary model but unstable in sensitivity analysis.
Original abstract
In this factor-informed exploratory CFA/SEM study, AI-literacy-related domains were treated as theoretically informed and empirically tested predictors of AI-TPACK readiness rather than as fully validated independent latent variables. Artificial intelligence (AI) is increasingly entering mathematics education, making it important to understand how preservice teachers become ready to integrate AI-supported tools pedagogically. This study examined AI-TPACK readiness among 130 preservice mathematics teachers at a South African public university. Exploratory factor analysis using polychoric correlations indicated that the AI-TPACK readiness items were essentially unidimensional; one weak design-confidence item was removed. The refined seven-item measurement model fitted better than the original eight-item specification, although discriminant-validity evidence for the broader AI-literacy-related domains was mixed. The primary gender-controlled latent SEM (sample n = 129) showed good approximate fit, χ 2 (602) = 789.92, p < .001, CFI = .981, TLI = .984, RMSEA = .049, SRMR = .082, and explained 53.0% of the variance in AI-TPACK readiness. Positive associations were observed for prior AI use, critical-ethical appraisal, and support/enablers. The support/enablers path had the largest standardised coefficient, but should be interpreted cautiously because the construct had marginal AVE and overlapped with information-source engagement. Year level was significant in the primary model but less stable in sensitivity analysis. Overall, the findings suggest that readiness was associated with direct AI experience and critical-ethical judgement, while the contribution of support/enablers remains provisional. The study contributes a cautious empirical account of AI-TPACK readiness in a Global South teacher education context.
- PaperJournal of Second Language Writing17 Jul 2026
Can GenAI deliver growth-oriented feedback? Evidence for enhancing learners’ growth mindset and adaptive motivation in second language writing
Yuan Yao, Mi Rong, Nigel Mantou Lou
A field experiment with 92 Chinese undergraduates compared GenAI-generated growth-oriented feedback to corrective feedback in L2 writing. Growth-oriented feedback enhanced students' growth mindset and maintained adaptive responses to mistakes, while corrective feedback led to a decline in adaptive responses. The change in growth mindset mediated effects on adaptive responses and writing performance.
Original abstract
Drawing on the mindsets theory, this study explores the use of generative artificial intelligence (GenAI) in producing growth-oriented feedback (i.e., feedback that aligns with the principles of growth mindset) and investigates its impact on students’ growth mindset, adaptive responses, and second language (L2) English writing performance. This field experiment was conducted at a university in China, involving 92 first- and second-year undergraduate students ( M age = 18.76, SD =.882; 80.4% males and 19.6% females). The participants were randomly assigned to an experimental group ( n = 49) or a control group ( n = 43) and completed pre- and post-questionnaires. Over the course of a semester, the participants completed three argumentative writing tasks. After each task, the experimental group received GenAI-generated growth-oriented feedback, whereas the control group received GenAI-generated corrective feedback. The results showed that growth-oriented feedback significantly enhanced students’ growth mindset and maintained their adaptive responses to mistakes at a relatively high level. In contrast, GenAI-generated corrective feedback led to a decline in adaptive responses to mistakes. Moreover, the change in growth mindset mediated the effect of feedback types on adaptive responses and writing performance. This study offers insights into the effectiveness of GenAI in generating growth-oriented feedback, highlighting its potential of GenAI feedback in L2 writing to go beyond error correction and foster motivational, behavioral, and academic development.
- PaperTESOL Quarterly17 Jul 2026
Tracing the Change in the Status of a Language Teacher in University Governance through English Medium Instruction
Ikuya Aizawa, Yuka Akiyama
A longitudinal case study traced how a top-down English medium instruction policy elevated a language teacher's status and participation in faculty governance at a Japanese graduate engineering school. The teacher gained collaboration with a senior content teacher, repositioning language support as a recognized component of engineering education.
Original abstract
This article examines how a top‐down English medium instruction (EMI) decision changed a language teacher's status and participation in faculty governance in a graduate engineering school at a national university in Japan. Using a longitudinal single‐case design, the study follows one language teacher across the period from policy announcement to early rollout of EMI implementation. Data comprise three milestone interviews, audio‐recorded research‐practice meetings, reflective journals and implementation documents. Two main findings emerged. First, EMI acted as a catalyst for cultural change by raising the status of the language teacher in a setting where language support had been dismissed as a distraction from “real” engineering. Second, EMI granted the language teacher access to collaborate with a senior content teacher, who served as a gatekeeper to institutional leadership. This collaboration resulted in the formal establishment of EMI support provision and repositioned language support as a recognized component of engineering education. We found that EMI transformed not only the teacher's participation in governance but also faculty‐wide decision‐making processes. The study demonstrates how process tracing can document the decision‐making mechanisms that contribute to institutional cultural change during a faculty‐level transition to EMI.
- PaperReCALL17 Jul 2026
The effect of high-immersion virtual reality environments on L2 learners’ socio-emotional variables: A meta-analysis
Yulia Khoruzhaya, Kara Moranski, Alexandra Neuenschwander, Nicole Ziegler
A meta-analysis of 18 studies found that high-immersion virtual reality (HiVR) significantly improved L2 learners' sense of presence, self-efficacy, motivation, and engagement, and reduced anxiety compared to low-immersion instruction, with an overall effect size of g = 0.50. Moderator analyses showed that effects varied by learners' proficiency level and target language but not by educational level, content, activities, or headset type.
Original abstract
The effect of high-immersion virtual reality (HiVR) on second language (L2) learners’ socio-emotional responses has gained increasing attention in recent years, though findings remain mixed. This study meta-analyzed 18 primary studies (31 effect sizes, n = 996) published between 2019 and 2024 to examine (a) the overall effect of HiVR on L2 learners’ socio-emotional variables as compared with low-immersion (2D) language instruction and (b) the extent to which this effect varies as a function of three groups of moderators, including specific socio-emotional variables as well as learner- and treatment-related characteristics. Regarding educational level, the included studies evaluated secondary and postsecondary L2 learners. Using a random-effects model, the resulting aggregate effect size across studies was g = 0.50. Subsequent moderator analyses showed that HiVR significantly increased L2 learners’ sense of presence, self-efficacy, motivation, and engagement, while reducing anxiety as compared to less immersive experiences. Additionally, variability in the effect of HiVR was explained by learners’ proficiency level and target language. No significant moderating effect was identified for educational level, HiVR content, HiVR learning activities, and head-mounted display type. These findings underscore HiVR’s potential as a powerful tool for improving socio-emotional outcomes in L2 learning.
- PaperComputers & Education17 Jul 2026
Designing effective feedback in video lectures using GenAI: The roles of feedback type and learner agency in error detection
Zhongling Pi, Yuxiang Cao, Yiliang Zhang, Yuxian Ma et al.
This study investigates how different types of GenAI-generated feedback and learner agency influence error detection in video lectures, aiming to design more effective feedback mechanisms.
- PaperApplied Linguistics16 Jul 2026
Shaping an enduring conversation on AI in applied linguistics
Ron Martinez, Glenn Martinez, Samantha Curle
This paper advocates for a sustained scholarly conversation on the role of artificial intelligence in applied linguistics, emphasizing the need for ongoing critical engagement.
- PaperETS Research Report Series16 Jul 2026
Second Language Writing Processes in the TOEFL iBT® Test: Examining the Write for an Academic Discussion Task Using Eye-Tracking, Keystroke Logging, and Stimulated Recall
Ching-Ni Hsieh, Renka Ohta
Using eye-tracking, keystroke logging, and stimulated recall, the study examined L2 writing processes in the TOEFL iBT Write for an Academic Discussion task. The integrated data revealed five core cognitive processes—prompt (re)reading, planning, formulation, monitoring, and revision—and showed that task design influences strategic behaviors and attention allocation, supporting construct validity.
Original abstract
The present study investigates the construct validity of the TOEFL iBT® writing test and focuses on the writing processes elicited by the Write for an Academic Discussion (WAD) task. Nineteen adult L2 English users completed the WAD task while their eye movements and keystrokes were recorded, followed by stimulated recall interviews to capture their writing strategies and thought processes. Eye-tracking data revealed sustained attention to the writing window, with frequent reference to the task question and example posts. Keystroke logging indicated substantial initial pauses and language-related revision behaviors. Qualitative analysis of stimulated recalls identified five core cognitive processes: prompt (re)reading, planning, formulation, monitoring, and revision. The integration of eye-tracking, keystroke logging, and stimulated recall demonstrates that the WAD task engages L2 writers in processes consistent with theoretical models of academic writing, supplying backing for the construct validity of the TOEFL iBT Writing test. The findings also reveal that task design features may influence L2 writing processes and shape writers’ strategic behaviors and attention allocation. Suggested citation: Hsieh, C.-N., & Ohta, R. (in press). Second language writing processes in the TOEFL iBT® test: Examining the Write for an Academic Discussion task using eye-tracking, keystroke logging, and stimulated recall (Research Report). ETS. https://doi.org/10.64634/tk1sd682
- PaperJournal of Second Language Writing16 Jul 2026
Conceptualizing voice in the age of AI: A response to Sandstead and Kibler (2025)
Paul Kei Matsuda, Xiao Tan
This response argues for a reconceptualization of voice in writing instruction in light of AI technologies, engaging with Sandstead and Kibler's (2025) work.
- PaperTESOL Quarterly16 Jul 2026
Fostering Fine‐Tuned Prompt Literacy for Multimodal Writing: A Systemic Functional Linguistics‐Informed Framework
Xiao Tan, Chaoran Wang, Ruonan Zhao
A pedagogical framework informed by Systemic Functional Linguistics (SFL) is introduced to develop fine-tuned prompt literacy for GenAI-assisted multimodal writing. Research on ESL students' photo essay creation shows that separating instruction on multimodal and GenAI literacies hinders intentional and rhetorically effective resource creation, while the SFL-based framework guides students in designing prompts using ideational, interpersonal, and textual metafunctions.
Original abstract
The rise of GenAI‐powered image generation tools presents new opportunities for English learners to engage in digital multimodal composing (DMC) in both creative and critical ways. At the same time, it calls for a reconceptualization of DMC pedagogy in language education. In this article, we build on the concept of fine‐tuned prompt literacy (Kang & Yi, 2023) in the context of GenAI‐assisted multimodal writing and introduce a pedagogical framework informed by Systemic Functional Linguistics (SFL) to support its development and implementation. We conceptualize fine‐tuned prompt literacy at the intersection of multimodal literacy and GenAI literacy, emphasizing the need for students to envision meaningful GenAI outcomes and engage with GenAI tools critically and strategically. Drawing on our research on ESL students' photo essay creation with GenAI image generators, we show that separating instruction on multimodal and GenAI literacies may hinder students' process of creating multimodal resources in an intentional and rhetorically effective way. To address this gap, we present an SFL‐informed framework that guides students in designing prompts based on the ideational, interpersonal, and textual metafunctions. Teaching observation and students' writing from a first‐year ESL writing course illustrates the framework's potential to enhance GenAI‐integrated DMC instruction. This brief report offers both theoretical insights and practical strategies for teaching multimodal writing in the era of GenAI.
- PaperLanguage Teaching Research16 Jul 2026
Comparing Teacher and Artificial Intelligence Scoring in Writing Assessment: A Generalizability Theory Analysis
Burak Asma
Middle school student essays were scored by Turkish teachers and AI tools, with and without rubrics. Generalizability theory analysis showed AI outperformed teachers in distinguishing individual differences and consistency, especially without rubrics. Teachers acknowledged biases and suggested rubrics and AI feedback for improvement.
Original abstract
This study examined the use of artificial intelligence tools, which have garnered significant attention in recent years, in the assessment and evaluation processes of language education. For this purpose, student essays were scored by Turkish middle school teachers and artificial intelligence tools both with and without the use of a rubric, and the findings were evaluated based on generalizability theory. Additionally, the research findings were shared with participants to gather qualitative data, which were analysed using the inductive thematic analysis method to support the research results. The findings revealed that in evaluations conducted without a rubric, teachers were limited in their ability to distinguish individual differences and demonstrated low scoring consistency. In contrast, artificial intelligence tools were more effective in distinguishing individual differences and exhibited high consistency. In evaluations conducted using a rubric, scoring consistency increased in both groups, although, as in the first evaluation, artificial intelligence tools demonstrated a higher level of consistency. Regarding the research findings, teachers expressed that individual biases, mood, and professional experiences influenced their scoring processes and emphasized the potential of rubrics and artificial intelligence-supported feedback systems for achieving more consistent results. Artificial intelligence tools, on the other hand, highlighted their independence from subjective factors but stressed the need for more diverse and generalizable datasets to further enhance their evaluation capacities.
- PaperLanguage Teaching Research16 Jul 2026
Transforming Traditional Translation Training: The Role of Policy, Practice, and Institutional Support in Student Skill Enhancement
Kang Gao, Wenjing Dong
This study examines the impact of English education policies on translation training across macro, meso, and micro levels. Using a mixed-methods design with interviews and surveys, the results show that translation training is significantly associated with policies and positively related to innovative models, career readiness, training quality, practical skills, and intercultural competence.
Original abstract
Globalization has revolutionized translation training, and English education policies have encouraged more practical, innovative, and, in this regard, oriented teaching approaches. This study investigates the multi-level effects of English education policies on translation training at the national (macro), institutional (meso), and classroom (micro) levels, with particular attention to how policy influences the adoption of innovative instructional models and students’ perceived learning-related outcomes. Using a mixed-methods design, qualitative data were gathered through semi-structured interviews with 78 respondents, including policymakers, administrators, instructors, and students, whereas quantitative data were collected through a survey of 436 translation students (undergraduate and postgraduate). Qualitative data were analyzed using NVivo, and SPSS and SmartPLS were used to analyze quantitative data. The results indicate that translation training (β = .373, p < .001) is significantly associated with English education policies. Moreover, translation training is positively related to innovative training models (β = .302), career readiness (β = .268), perceived training quality (β = .318), practical skill acquisition (β = .292), and intercultural competence (β = .285), all at p < .001. These findings illustrate the focal point of translation training as one of the critical mechanisms through which policy is connected to various perceived academic and professional outcomes. University cooperation and pedagogical development at the institutional and classroom levels are also associated with perceived quality and applicability of translation training programs. Its uniqueness lies in that it combines several dimensions of outcomes in one analytical framework and provides a clear understanding of policy–practice alignment. The results suggest that translating students need effective policy–practice integration to be associated with students’ perceived preparedness for global professional contexts.