prompt-engineering
Filtering by topic prompt-engineering(6)Clear all filters
- PaperComputers & Education10 Jul 2026
Educational prompt engineering self-efficacy scale (Ed-PESS): Development and psychometric validation
Fatih Karataş, Recep GÜR, Barış Eriçok, Fatma BAŞARIR et al.
Develops and psychometrically validates the Educational Prompt Engineering Self-Efficacy Scale (Ed-PESS) for measuring educators' self-efficacy in prompt engineering within educational contexts.
- PaperarXiv — AI in Education (cs.CY)7 Jul 2026
Say What? Examining Text and Voice Input Modalities for Prompt-Based Programming in Computing Education
Kaitlin Riegel, Yan Cathy Hua, Paul Denny, Victor-Alexandru Pădurean et al.
Large language models are increasingly used in computing education, but most research focuses on text-based interactions. This study examines how introductory programming students perform with text versus voice input when crafting prompts to generate code, finding that text prompts succeed more on first attempts, though editing transcribed voice prompts eliminates the difference.
Original abstract
Large language models (LLMs) are increasingly integrated into computing education, yet nearly all prior research has focused on text-based interactions. As voice-enabled interfaces become more capable and more common, there is growing interest in understanding how voice input might shape students' use of LLM-powered tools. In this exploratory study, we investigated how introductory programming students interact with Prompt Problems, which are programming tasks that require crafting natural-language prompts to generate correct code. Students (N = 919) solved a series of Prompt Problems with the freedom to select or switch between text and voice input modalities. We collected their prompt submissions as well as post-activity survey responses, then analysed differences in prompt accuracy, persistence, and perspectives by modality. For two of the three problems, we found that students who typed their prompts using text were more likely to have those prompts succeed on the first attempt than students who submitted unedited voice prompts. There was no difference in success rate if students edited their transcribed voice prompts before submission. Across the problems, we found evidence that students who tried voice prompting varied in their usage of modality - perhaps indicating a complementary, or non-preferential approach. However, most students only tried and reported preferring text. Our qualitative analysis revealed how students' perceived the roles of voice and text input in shaping their problem-solving process, as well as the reported drawbacks and advantages of each modality. We discuss implications for future multimodal tools and instructional design in computing education.
- PaperarXiv — AI in Education (cs.CY)29 Jun 2026
Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers
Keith Tran, Samiha Marwan, Thomas Price
Evaluated a 45-minute lesson with guided practice on prompt-based programming using LLMs against a standard CS lab activity. Engineering students in the experimental group showed modest performance gains and significantly greater increases in prompting self-efficacy, suggesting brief interventions can improve prompt specification skills.
Original abstract
Prompt-based programming, a new modality enabled by large language models (LLMs), allows users to express computational goals through natural language rather than traditional code. While this approach lowers barriers to entry, especially for non-CS learners, it does not eliminate the need for foundational CS skills. Learners often struggle to communicate their intent clearly to LLMs, resulting in vague or underspecified prompts. Prior work has documented the need for explicit prompting for both CS and non-CS learners. However, it remains less clear how such instruction can fit into busy classrooms or how much time is needed to produce meaningful gains. In this paper, we evaluated a 45-minute prompt-based programming intervention, consisting of a lesson with guided practice, against a business-as-usual CS lab activity (code tracing) of equal length, representing a class without prompt-focused instruction. We conducted a randomized controlled study with 55 engineering students. We found that students in the experimental condition improved more on average (though not significantly more) from pre- to post-test than the control group (+10.8 vs +1.1 percentage points) and showed significantly greater average gains in prompting self-efficacy (+35.4 vs +21.9 percentage points). Our results suggest it is likely that a brief intervention can improve learners' ability to specify computational goals to LLMs. However, the effect was modest, suggesting that prompting skills may require more time and practice to develop. We provide a lightweight lesson that requires no prior CS background and can be readily dropped into existing courses.
- PaperComputers and Education: Artificial Intelligence21 Jun 2026
Students’ multimodal prompting practices as epistemic work in AI literacy development
Sylvana Sofkova Hashemi
Postgraduate students engaged in collaborative multimodal prompting tasks with a GenAI tool, revealing a continuum of prompting strategies from basic input-output to strategic, iterative, and dialogic practices. Prompting emerges as an epistemic practice for critically interpreting and negotiating AI outputs, fostering AI literacy but with underdeveloped ethical dimensions.
Original abstract
: As generative artificial intelligence (GenAI) rapidly transforms higher education, critical questions arise about how students engage with these open-ended tools and the implications for learning. This study provides empirical insight into this research gap investigating (1) the prompting strategies students develop when interacting with a university-provided GenAI tool and (2) how engagement in prompt engineering activities shapes their understanding of GenAI and AI literacy. Data were collected in an exploratory workshop with 28 postgraduate students engaged in collaborative multimodal prompting tasks, including the creation of short stories or poems and corresponding images. Students’ self-documented prompting histories and reflections were analysed qualitatively using reflexive thematic analysis, guided by frameworks for prompting methods and AI literacy. The findings show that students’ prompting strategies vary along a continuum from basic input-output use to strategic, iterative, and dialogic practices. Prompting emerges as a central epistemic practice through which students critically interpret, refine, and negotiate AI-generated outputs. Multimodal engagement exposes challenges in translating abstract meaning into machine-readable prompts, fostering awareness of system limitations, bias, and the need to actively construct coherence across modalities. While students demonstrate developing competence in evaluation and creation, ethical dimensions of AI literacy remain underdeveloped. The findings provide empirical insight into how AI literacy develops through hands-on engagement with GenAI, positioning prompting as an epistemic practice through which students learn to interpret, negotiate, and guide AI-generated outputs, while highlighting the value of iterative, reflective, and multimodal learning designs that foster critical, strategic, and responsible engagement with AI.
- PaperLanguage Testing25 Mar 2026
™ChatGPT for automated writing evaluation: Scoring and feedback across prompt conditions
Yewon Lee, Myunghwan Hwang
Six ChatGPT models with different prompt configurations were tested against human raters on 60 EFL writing samples. Prompt design significantly influenced scoring consistency and severity, with Chain-of-Thought and Fill-in-the-blank prompts yielding higher reliability. Learners perceived the feedback positively, but reasoning-intensive domains still required human oversight.
Original abstract
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
- PaperReCALL29 Dec 2025
Impact of prompt sophistication on ChatGPT’s output for automated written corrective feedback
Na Luo, Yifan Wang, Zhe (Victor) Zhang, Yile Zhou et al.
Compared ChatGPT's automated written corrective feedback (AWCF) accuracy across generic and domain-specific prompts against Grammarly. Found that domain-specific prompts, especially one-shot, significantly improved error detection, with zero-shot matching Grammarly and one-shot surpassing it. However, even the most sophisticated prompt still showed limitations compared to Grammarly.
Original abstract
The emergence of large language models, exemplified by ChatGPT, has garnered growing attention for their potential to generate feedback in second language writing, particularly automated written corrective feedback (AWCF). In this study, we examined how prompt design – a generic prompt and two domain-specific prompts (zero-shot and one-shot) enriched with comprehensive domain knowledge about written corrective feedback (WCF) – influences ChatGPT’s ability to provide AWCF. The accuracy and coverage of ChatGPT’s feedback across these three prompts were benchmarked against Grammarly, a widely used traditional automated writing evaluation (AWE) tool. We find that ChatGPT’s ability in flagging language errors grew considerably with prompt sophistication driven by the integration of domain-specific knowledge and examples. While the generic prompt resulted in substantially lower performance than Grammarly, the zero-shot prompt achieved comparable results to it and the one-shot prompt surpassed it considerably in error detection. Notably, the most pronounced improvement in ChatGPT’s performance was observed in its detection of frequent error categories, including those of word choice or expression, direct translation, sentence structure and pronoun. Nonetheless, even with the most sophisticated prompt, ChatGPT still displayed certain limitations when compared to Grammarly. Our study has both theoretical and practical implications. Theoretically, it lends empirical evidence to Knoth et al .’s (2024) proposition to separate domain-specific AI literacy from generic AI literacy. Practically, it sheds light on the pedagogical application and technical development of AWE systems.