prompt-engineering
Filtering by topic prompt-engineering(5)Clear all filters
- PaperComputers & Education10 Jul 2026
Educational prompt engineering self-efficacy scale (Ed-PESS): Development and psychometric validation
Fatih Karataş, Recep GÜR, Barış Eriçok, Fatma BAŞARIR et al.
A new scale called the Educational Prompt Engineering Self-Efficacy Scale (Ed-PESS) was developed and psychometrically validated to measure educators' confidence in using prompt engineering for educational purposes.
- PaperarXiv — AI in Education (cs.CY)29 Jun 2026
Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers
Keith Tran, Samiha Marwan, Thomas Price
A 45-minute lesson on prompt-based programming with LLMs was evaluated against a traditional lab activity. Students receiving the prompt-focused instruction showed greater gains in prompting self-efficacy and modest improvements in programming performance. The results suggest that brief interventions can help but more practice is likely needed.
Original abstract
Prompt-based programming, a new modality enabled by large language models (LLMs), allows users to express computational goals through natural language rather than traditional code. While this approach lowers barriers to entry, especially for non-CS learners, it does not eliminate the need for foundational CS skills. Learners often struggle to communicate their intent clearly to LLMs, resulting in vague or underspecified prompts. Prior work has documented the need for explicit prompting for both CS and non-CS learners. However, it remains less clear how such instruction can fit into busy classrooms or how much time is needed to produce meaningful gains. In this paper, we evaluated a 45-minute prompt-based programming intervention, consisting of a lesson with guided practice, against a business-as-usual CS lab activity (code tracing) of equal length, representing a class without prompt-focused instruction. We conducted a randomized controlled study with 55 engineering students. We found that students in the experimental condition improved more on average (though not significantly more) from pre- to post-test than the control group (+10.8 vs +1.1 percentage points) and showed significantly greater average gains in prompting self-efficacy (+35.4 vs +21.9 percentage points). Our results suggest it is likely that a brief intervention can improve learners' ability to specify computational goals to LLMs. However, the effect was modest, suggesting that prompting skills may require more time and practice to develop. We provide a lightweight lesson that requires no prior CS background and can be readily dropped into existing courses.
- PaperComputers and Education: Artificial Intelligence21 Jun 2026
Students’ multimodal prompting practices as epistemic work in AI literacy development
Sylvana Sofkova Hashemi
Investigates prompting strategies of 28 postgraduate students using a generative AI tool in collaborative multimodal tasks, finding strategies ranging from basic input-output to strategic, iterative, and dialogic practices. Prompting emerges as an epistemic practice for AI literacy, fostering critical interpretation and awareness of system limitations, while ethical dimensions remain underdeveloped. The study highlights the value of iterative, reflective, and multimodal learning designs for fostering critical and strategic engagement with AI.
Original abstract
: As generative artificial intelligence (GenAI) rapidly transforms higher education, critical questions arise about how students engage with these open-ended tools and the implications for learning. This study provides empirical insight into this research gap investigating (1) the prompting strategies students develop when interacting with a university-provided GenAI tool and (2) how engagement in prompt engineering activities shapes their understanding of GenAI and AI literacy. Data were collected in an exploratory workshop with 28 postgraduate students engaged in collaborative multimodal prompting tasks, including the creation of short stories or poems and corresponding images. Students’ self-documented prompting histories and reflections were analysed qualitatively using reflexive thematic analysis, guided by frameworks for prompting methods and AI literacy. The findings show that students’ prompting strategies vary along a continuum from basic input-output use to strategic, iterative, and dialogic practices. Prompting emerges as a central epistemic practice through which students critically interpret, refine, and negotiate AI-generated outputs. Multimodal engagement exposes challenges in translating abstract meaning into machine-readable prompts, fostering awareness of system limitations, bias, and the need to actively construct coherence across modalities. While students demonstrate developing competence in evaluation and creation, ethical dimensions of AI literacy remain underdeveloped. The findings provide empirical insight into how AI literacy develops through hands-on engagement with GenAI, positioning prompting as an epistemic practice through which students learn to interpret, negotiate, and guide AI-generated outputs, while highlighting the value of iterative, reflective, and multimodal learning designs that foster critical, strategic, and responsible engagement with AI.
- PaperLanguage Testing25 Mar 2026
™ChatGPT for automated writing evaluation: Scoring and feedback across prompt conditions
Yewon Lee, Myunghwan Hwang
ChatGPT's performance as an Automated Writing Evaluation (AWE) system was tested by comparing its scoring with human raters on 60 EFL essays from Korean university students. Prompt design significantly affected scoring consistency, with Chain-of-Thought reasoning combined with Fill-in-the-blank scaffolding yielding higher reliability, though domain-level bias remained uncontrolled. Learners generally perceived ChatGPT's feedback positively, suggesting prompt-calibrated AWE can serve as a supplementary tool for writing assessment.
Original abstract
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
- PaperReCALL29 Dec 2025
Impact of prompt sophistication on ChatGPT’s output for automated written corrective feedback
Na Luo, Yifan Wang, Zhe (Victor) Zhang, Yile Zhou et al.
The study examined how prompt design (generic vs. domain-specific zero-shot and one-shot) affects ChatGPT's accuracy and coverage in providing automated written corrective feedback (AWCF), benchmarked against Grammarly. Domain-specific prompts, especially one-shot, significantly improved error detection, surpassing Grammarly in some categories. However, ChatGPT still showed limitations compared to Grammarly.
Original abstract
The emergence of large language models, exemplified by ChatGPT, has garnered growing attention for their potential to generate feedback in second language writing, particularly automated written corrective feedback (AWCF). In this study, we examined how prompt design – a generic prompt and two domain-specific prompts (zero-shot and one-shot) enriched with comprehensive domain knowledge about written corrective feedback (WCF) – influences ChatGPT’s ability to provide AWCF. The accuracy and coverage of ChatGPT’s feedback across these three prompts were benchmarked against Grammarly, a widely used traditional automated writing evaluation (AWE) tool. We find that ChatGPT’s ability in flagging language errors grew considerably with prompt sophistication driven by the integration of domain-specific knowledge and examples. While the generic prompt resulted in substantially lower performance than Grammarly, the zero-shot prompt achieved comparable results to it and the one-shot prompt surpassed it considerably in error detection. Notably, the most pronounced improvement in ChatGPT’s performance was observed in its detection of frequent error categories, including those of word choice or expression, direct translation, sentence structure and pronoun. Nonetheless, even with the most sophisticated prompt, ChatGPT still displayed certain limitations when compared to Grammarly. Our study has both theoretical and practical implications. Theoretically, it lends empirical evidence to Knoth et al .’s (2024) proposition to separate domain-specific AI literacy from generic AI literacy. Practically, it sheds light on the pedagogical application and technical development of AWE systems.