assessment-design
Filtering by topic assessment-design(6)Clear all filters
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning
Ahmad D. Suleiman, Daqing Hou, Maliha Noushin Raida
This study introduces design problems (DPs)—concise, scenario-based prompts—to assess higher-order thinking in project-based learning. Surveys of 31 instructors and evaluations of 80 LLM-generated DPs showed that instructors value DPs but find creation effort a barrier, while LLMs can generate high-quality prompts. Student performance data indicated that DPs capture distinct aspects of higher-order thinking, with negligible correlation to traditional project grades.
Original abstract
Project-based learning (PjBL) is common in computing education, but traditional assessments of PjBL often fail to capture higher-order thinking (HOT), especially in transfer contexts. This study introduces "design problems" (DPs): concise, scenario-based prompts that require applying project concepts in new situations, to address this gap. We examined instructor perceptions, the ability of large language models (LLMs) to generate DPs, and student experiences. Surveys of 31 instructors, evaluation of 80 LLM-generated DPs, and student performance data showed that while instructors value DPs, creation effort is a barrier. LLMs helped by producing high-quality prompts with strong expert agreement. Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs may capture distinct aspects of HOT. Keystroke data also suggested deeper cognitive engagement of students through planning and revision behaviors. Overall, DPs appear to be a useful complement to traditional assessments, especially in situations where AI use or collaboration may undermine individual learning.
- PaperETS Research Report Series10 Jul 2026
Identifying High-Leverage Practices for Guiding the Development of Teaching Assessments
Geoffrey Phelps, Heather Howell, Jamie Mikeska, Caroline Wylie
This report identifies eight high-leverage practices (HLPs) for teaching, drawn from an existing framework of 34 HLPs, to guide the development of a next generation of teaching assessments. The methods for selecting these HLPs are described, and initial recommendations for appropriate assessment methods for each HLP are provided.
Original abstract
This report identifies eight high-leverage practices (HLPs) for teaching that are intended for use in guiding the development of a next generation of teaching assessments. The identified HLPs are drawn from an existing framework that provides empirical and theoretical research backing for 34 HLPs that make up the work of teaching. The report describes the methods used to select a subset of eight HLPs and provides initial recommendations for the assessment methods that are most appropriate for each of these eight HLPs.
- PaperAssessment & Evaluation in Higher Education4 Jul 2026
From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich higher education
Vangelis Tsiligiris
This systematic conceptual review synthesizes foundational and contemporary literature on authentic assessment in higher education under AI-rich conditions, leading to a six-dimension framework. The authors argue that authentic assessment should not be reduced to workplace simulation or treated primarily as a response to academic misconduct; instead, it is a multidimensional design orientation spanning contextual fidelity, cognitive demand, process transparency, student agency, inclusivity, and AI-aware validity. The key contribution is distinguishing authentic products from authenticated processes, emphasizing that validity under generative AI requires architectures that make human judgement and responsibility visible.
Original abstract
This article presents a systematic conceptual review of authentic assessment in higher education and develops a six-dimension framework for assessment design in digitally mediated and AI-rich conditions. Drawing on a retained corpus of 37 substantive sources, it synthesises foundational and contemporary literature on task fidelity, evaluative judgement, process evidence, inclusion, and AI-mediated validity. The synthesis shows that authentic assessment should not be reduced to workplace simulation or treated primarily as a response to academic misconduct. It is better understood as a multidimensional design orientation spanning contextual fidelity and consequential relevance, cognitive demand and evaluative judgement, process transparency and integrity, student agency and bounded choice, inclusivity and representational fairness, and AI-aware validity and ethical practice. The article’s main contribution is to distinguish authentic products from authenticated processes. It argues that assessment validity under generative AI depends not only on realistic outputs, but on architectures that make human judgement, verification, and responsibility visible. The framework offers review questions that support module-level and programme-level redesign by linking authenticity, evidence, validity, and accountable student judgement.
- PaperAssessment & Evaluation in Higher Education3 Jul 2026
The potential of authentic assessment in literary studies pedagogy and the problem with the ‘real world’
Sabine Kildea, Isobel Lavers, Hannah Upton, Claire Hansen
This article explores authentic assessment in literary studies, implementing a pilot 'Great Writers Festival' assessment that combined creative and critical options, groupwork, peer review, and a public event. It critiques the concept of 'real world' applicability in authentic assessment, arguing for a more nuanced understanding in humanities contexts.
Original abstract
This article explores the potential of authentic assessment in literary studies education while offering a critique of the understanding of authentic assessment as a means of fostering student experience in ‘real world’ contexts. We consider authentic assessment as it pertains to literary studies and the humanities and analyse the implementation of a pilot authentic assessment design at an Australian university. The pilot assessment, titled ‘The Great Writers Festival’, comprised creative and traditional critical assessment options, a blend of groupwork and individual contributions, peer review and self-evaluation, and culminated in a public event in which all students presented their final projects. Our article takes a twofold approach, in that while examining the implementation of an authentic assessment model in literary studies, we also interrogate the concept of ‘real world’ applicability in authentic assessment literature and practice. The essay analyses student assessment, self-evaluation and survey responses and evaluates the implementation of authentic assessment in a humanities discipline through a framework structured by five key traits: critical thinking; interpersonal skills; engagement; ontological knowledge and social impact.
- PaperAssessment & Evaluation in Higher Education3 Jul 2026
AI-resistant and AI-resilient assessment in higher education: a systematic review of validity-grounded strategies, institutional frameworks, and equity implications
Asrat Genet Amnie
This systematic review differentiates between AI-resistant and AI-resilient assessment strategies in higher education. It synthesizes 47 peer-reviewed studies and 34 policy documents to propose a Six-Layer AI-Resilient Assessment Stack grounded in validity theory. The review finds that AI-detection tools are insufficient and highlights challenges in equity, reliability, and faculty development.
Original abstract
The emergence of large language models has precipitated a fundamental disruption to higher education assessment. Conventional instruments are susceptible to AI-assisted completion, threatening construct validity across disciplines and institutional contexts. When submitted work reflects AI capability rather than student competency, the inferential chain from performance to qualification is invalidated – a problem of design rather than detection. This systematic review makes two contributions. First, it theorises and grounds the distinction between AI-resistant assessment, which seeks to prevent or impede AI access through containment, and AI-resilient assessment, which assesses genuine human cognitive performance of intrinsic educational value regardless of the AI tools that exist. Second, it synthesises peer-reviewed empirical literature, policy documentation, and sector guidance published between 2022 and 2025. A PRISMA 2020-compliant search of five databases was conducted; two independent reviewers screened all records with substantial inter-rater agreement. Forty-seven peer-reviewed studies, 23 policy documents, and 11 sector reports were included. Six evidence-based assessment categories were identified, taxonomised, and evaluated against a structured adversarial threat taxonomy. AI-detection tools are structurally insufficient as primary responses. A Six-Layer AI-Resilient Assessment Stack, grounded in multi-trait multi-method validation logic, is proposed as an integrated institutional framework. Persistent challenges include equity, reliability, faculty development, and data protection.
- PaperarXiv — AI in Education (cs.CY)29 Jun 2026
Demystify, Use, Reflect, Assess (DURA): An Experience Report on LLM Integration in CS2
Margaret Ellis, Nikitha Donekal Chandrashekar, Sehrish Basir Nizamani, Mohammed Farghally et al.
The DURA framework (Demystify, Use, Reflect, Assess) guides the integration of Large Language Models into a CS2 course, including demystification, guided use with attribution, reflective activities, and adjusted assessments. Students reported positive engagement with LLMs while still valuing traditional support, and perceptions of instructor care improved.
Original abstract
Student access to Large Language Models (LLMs) is reshaping learning behaviors; at the same time students are entering the workforce where effective LLM use is becoming an expected skill. In this Experience Report we share our DURA framework (Demystify-Use-Reflect-Assess) and materials we used to restructure our CS2 course to allow the use of LLMs. We first demystified LLMs, then provided guidance on use with required attribution. We also added reflections related to LLM use at three points throughout the semester to encourage student meta-cognition around LLM use. We increased the value of proctored assessments in tandem with allowing retakes and including questions that explicitly assess skills from programming assignments. Students reported using LLMs for clarifying course concepts, debugging, understanding assignment guidelines, and determining test cases, but also still sought assistance via office hours and TAs, monitored Piazza, and reviewing course content. Students articulated thoughtful and strategic approaches to LLM use and also valued the instructional content and guidance from course staff. Student use of office hours increased slightly this semester and student perceptions that the instructor cares about them and their learning improved.