programming-education
Filtering by topic programming-education(5)Clear all filters
- PaperarXiv — AI in Education (cs.CY)17 Jul 2026
EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha, Jungpil Shin
EduGuard, a safe retrieval-augmented generation tutoring framework for introductory programming, integrates query understanding, instructor-approved course retrieval, pedagogical strategy selection, rubric-aware generation, claim-level verification, and overreliance control. On a 600-query benchmark and a pilot study with 10 undergraduates, it achieved 90.1% correctness, 4.9% hallucination, and reduced overreliance from 38% to 17% compared to GPT-4o-mini Tutor. The results demonstrate that safe GenAI tutoring requires explicit pedagogical control and evidence verification beyond retrieval or prompting alone.
Original abstract
Generative AI (GenAI) is increasingly used by students for programming explanation, debugging, and assignment support. Yet unrestricted large language model (LLM) tutors can hallucinate, contradict course policy, reveal complete solutions, and foster passive dependence. This paper presents EduGuard, a safe retrieval-augmented generation (RAG) tutoring framework for introductory programming. EduGuard integrates query understanding, instructor-approved course retrieval, pedagogical strategy selection, rubric-aware generation, claim-level verification, and overreliance control. To make evaluation provenance explicit, we construct BILearn-CS, a 600-query instructor-authored, TA-validated benchmark spanning concept questions, debugging cases, misconceptions, assignment-support requests, code-mixed Bangla-English queries, and adversarial direct-answer prompts. Moving beyond a synthetic-only benchmark, we further evaluate on a 150-query public CS50-style course-forum set and run a small controlled pilot with 10 undergraduates using a counterbalanced pre-test/post-test design. Using Meta-Llama-3.1-8B-Instruct as the primary generator, hybrid FAISS/BM25 retrieval, and DeBERTa-v3-large-MNLI as an architecturally separate verifier, EduGuard is compared against strong baselines: GPT-4o-mini Tutor, Llama Socratic Tutor, LPITutor-style RAG, RAG with rubric prompting, and RAG with same-model self-checking. On BILearn-CS, EduGuard attains the best correctness (90.1%), grounding (89.4%), and rubric alignment (90.8%), with the lowest hallucination (4.9%) and direct-answer leakage (9.8%). In the pilot, it raises immediate post-test accuracy from 68.4% to 81.2% and cuts overreliance from 38.0% to 17.0% relative to GPT-4o-mini Tutor. These results suggest safe GenAI tutoring requires not only retrieval or strong prompting, but explicit pedagogical control, evidence verification, and deployment safeguards.
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis
Adrian-Marius Dumitran, Iulia-Maria Popescu
A 15-nation comparative analysis finds that many students complete secondary education without formal programming exposure, and among those who do, a 'Syntax Ceiling' limits algorithmic depth: Python is widespread but C++ remains in elite STEM tracks. Governance structures and high-stakes exams, rather than curriculum content alone, drive these inequities, undermining the goal of universal AI literacy.
Original abstract
The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by governance decisions made largely with reference to the specialist one. This paper presents a comparative analysis of curricula and examination frameworks across fifteen countries, identifying two structural challenges. First, in several systems a significant portion of students completes secondary education without any formal programming exposure. Second, among those who do receive CS education, a \emph{Syntax Ceiling} emerges: Python-based instruction reaches most students, while the algorithmic depth associated with C++ remains concentrated in elite STEM tracks. Drawing on reform cases spanning centralised mandates (France, China, Japan), assessment-driven systems (Poland, Romania, South Korea), and recent universal reforms (Switzerland, Kazakhstan), we show that governance structures and high-stakes examinations are the primary drivers of both challenges -- and that specialist and general-track language choices are rarely independent, linked through shared teacher pipelines that curriculum policy seldom acknowledges. Achieving genuine AI literacy for all requires confronting not just curriculum content, but the access architectures and resource constraints that determine who receives it -- and at what depth.
- PaperarXiv — AI in Education (cs.CY)12 Jul 2026
Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications
Nasser Giacaman, Valerio Terragni, Paul Denny, Viraj Kumar
Analyzed four years of undergraduate programming submissions to understand how students write natural-language comments as specifications for AI code generation. Introduced a taxonomy of comment types, code expression levels, and code constructs, finding that students predominantly wrote 'What' comments and shifted toward 'How' comments for procedural tasks, focusing more on verifying generated code than iterating on comments.
Original abstract
As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and rely on these tools to generate code, shifting emphasis from code writing to specification. Yet little is known about the comments students write as specifications in AI-assisted programming tasks. We analyze a four-year dataset of undergraduate programming submissions and reflections from tasks in which students wrote comments to guide code generation and refined solutions using test-case feedback. We introduce a taxonomy spanning three dimensions: comment type, code expression level, and code construct. Using automated classification, we examine how these dimensions vary across attempts and how students describe the process in their reflections. Our findings show that students mostly wrote natural-language What comments, shifted toward How comments for more procedural constructs, and focused more on verifying generated code than on repeatedly rewriting comments.
- PaperarXiv — AI in Education (cs.CY)9 Jul 2026
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
Yi Zhang, Julia Rayz
A Bloom-aligned framework is introduced to measure educational control in LLMs, focusing on shifting cognitive demand in programming tasks. Tests on 2,520 tasks using Qwen3 models reveal a directional asymmetry: models reliably increase cognitive demand but struggle to lower it. Strong task execution does not guarantee Bloom-aligned educational control.
Original abstract
We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its cognitive demand toward specified learning objectives. We apply this framework to programming tasks in computer science education to study the gap between solving tasks and adapting them for learners. Using revised Bloom's Taxonomy as an operational scale of cognitive demand, we evaluate two intervention settings: general difficulty control, where models are asked to make tasks harder or easier, and Bloom's control, where models are asked to target higher or lower Bloom's levels. We evaluate a matched Qwen3-Next model pair, comparing Qwen3-Next-80B-A3B-Instruct with Qwen3-Coder-Next across 2,520 tasks from three benchmarks. The framework reveals a robust directional asymmetry: both models reliably increase cognitive demand, but struggle to lower it. We further characterize these outcomes with semantic-delta clustering and layer-wise Fisher's Discriminant Ratio probing. Within this controlled comparison, the general model shows clearer middle-layer separability for both general difficulty and Bloom-control contrasts, whereas the coder model shows weaker separability for general difficulty and a deeper peak for Bloom-control contrasts. These results show that strong execution performance does not automatically entail Bloom-aligned educational control.
- PaperarXiv — AI in Education (cs.CY)29 Jun 2026
Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers
Keith Tran, Samiha Marwan, Thomas Price
A 45-minute prompt-based programming intervention with guided practice was evaluated against a control activity. Engineering students in the experimental group showed greater gains in prompting self-efficacy and slight improvement in ability to specify computational goals to LLMs. The results suggest that even a brief intervention can modestly improve prompting skills, but more time may be needed for significant gains.
Original abstract
Prompt-based programming, a new modality enabled by large language models (LLMs), allows users to express computational goals through natural language rather than traditional code. While this approach lowers barriers to entry, especially for non-CS learners, it does not eliminate the need for foundational CS skills. Learners often struggle to communicate their intent clearly to LLMs, resulting in vague or underspecified prompts. Prior work has documented the need for explicit prompting for both CS and non-CS learners. However, it remains less clear how such instruction can fit into busy classrooms or how much time is needed to produce meaningful gains. In this paper, we evaluated a 45-minute prompt-based programming intervention, consisting of a lesson with guided practice, against a business-as-usual CS lab activity (code tracing) of equal length, representing a class without prompt-focused instruction. We conducted a randomized controlled study with 55 engineering students. We found that students in the experimental condition improved more on average (though not significantly more) from pre- to post-test than the control group (+10.8 vs +1.1 percentage points) and showed significantly greater average gains in prompting self-efficacy (+35.4 vs +21.9 percentage points). Our results suggest it is likely that a brief intervention can improve learners' ability to specify computational goals to LLMs. However, the effect was modest, suggesting that prompting skills may require more time and practice to develop. We provide a lightweight lesson that requires no prior CS background and can be readily dropped into existing courses.