assessment-design
Filtering by topic assessment-design(9)Clear all filters
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning
Ahmad D. Suleiman, Daqing Hou, Maliha Noushin Raida
This study introduces design problems (DPs)—concise, scenario-based prompts—to assess higher-order thinking in project-based learning. Surveys of 31 instructors and evaluations of 80 LLM-generated DPs showed that instructors value DPs but find creation effort a barrier, while LLMs can generate high-quality prompts. Student performance data indicated that DPs capture distinct aspects of higher-order thinking, with negligible correlation to traditional project grades.
Original abstract
Project-based learning (PjBL) is common in computing education, but traditional assessments of PjBL often fail to capture higher-order thinking (HOT), especially in transfer contexts. This study introduces "design problems" (DPs): concise, scenario-based prompts that require applying project concepts in new situations, to address this gap. We examined instructor perceptions, the ability of large language models (LLMs) to generate DPs, and student experiences. Surveys of 31 instructors, evaluation of 80 LLM-generated DPs, and student performance data showed that while instructors value DPs, creation effort is a barrier. LLMs helped by producing high-quality prompts with strong expert agreement. Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs may capture distinct aspects of HOT. Keystroke data also suggested deeper cognitive engagement of students through planning and revision behaviors. Overall, DPs appear to be a useful complement to traditional assessments, especially in situations where AI use or collaboration may undermine individual learning.
- PaperETS Research Report Series10 Jul 2026
Identifying High-Leverage Practices for Guiding the Development of Teaching Assessments
Geoffrey Phelps, Heather Howell, Jamie Mikeska, Caroline Wylie
This report identifies eight high-leverage practices (HLPs) for teaching, drawn from an existing framework of 34 HLPs, to guide the development of a next generation of teaching assessments. The methods for selecting these HLPs are described, and initial recommendations for appropriate assessment methods for each HLP are provided.
Original abstract
This report identifies eight high-leverage practices (HLPs) for teaching that are intended for use in guiding the development of a next generation of teaching assessments. The identified HLPs are drawn from an existing framework that provides empirical and theoretical research backing for 34 HLPs that make up the work of teaching. The report describes the methods used to select a subset of eight HLPs and provides initial recommendations for the assessment methods that are most appropriate for each of these eight HLPs.
- PaperAssessment & Evaluation in Higher Education4 Jul 2026
From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich higher education
Vangelis Tsiligiris
This systematic conceptual review synthesizes foundational and contemporary literature on authentic assessment in higher education under AI-rich conditions, leading to a six-dimension framework. The authors argue that authentic assessment should not be reduced to workplace simulation or treated primarily as a response to academic misconduct; instead, it is a multidimensional design orientation spanning contextual fidelity, cognitive demand, process transparency, student agency, inclusivity, and AI-aware validity. The key contribution is distinguishing authentic products from authenticated processes, emphasizing that validity under generative AI requires architectures that make human judgement and responsibility visible.
Original abstract
This article presents a systematic conceptual review of authentic assessment in higher education and develops a six-dimension framework for assessment design in digitally mediated and AI-rich conditions. Drawing on a retained corpus of 37 substantive sources, it synthesises foundational and contemporary literature on task fidelity, evaluative judgement, process evidence, inclusion, and AI-mediated validity. The synthesis shows that authentic assessment should not be reduced to workplace simulation or treated primarily as a response to academic misconduct. It is better understood as a multidimensional design orientation spanning contextual fidelity and consequential relevance, cognitive demand and evaluative judgement, process transparency and integrity, student agency and bounded choice, inclusivity and representational fairness, and AI-aware validity and ethical practice. The article’s main contribution is to distinguish authentic products from authenticated processes. It argues that assessment validity under generative AI depends not only on realistic outputs, but on architectures that make human judgement, verification, and responsibility visible. The framework offers review questions that support module-level and programme-level redesign by linking authenticity, evidence, validity, and accountable student judgement.
- PaperAssessment & Evaluation in Higher Education3 Jul 2026
The potential of authentic assessment in literary studies pedagogy and the problem with the ‘real world’
Sabine Kildea, Isobel Lavers, Hannah Upton, Claire Hansen
This article explores authentic assessment in literary studies, implementing a pilot 'Great Writers Festival' assessment that combined creative and critical options, groupwork, peer review, and a public event. It critiques the concept of 'real world' applicability in authentic assessment, arguing for a more nuanced understanding in humanities contexts.
Original abstract
This article explores the potential of authentic assessment in literary studies education while offering a critique of the understanding of authentic assessment as a means of fostering student experience in ‘real world’ contexts. We consider authentic assessment as it pertains to literary studies and the humanities and analyse the implementation of a pilot authentic assessment design at an Australian university. The pilot assessment, titled ‘The Great Writers Festival’, comprised creative and traditional critical assessment options, a blend of groupwork and individual contributions, peer review and self-evaluation, and culminated in a public event in which all students presented their final projects. Our article takes a twofold approach, in that while examining the implementation of an authentic assessment model in literary studies, we also interrogate the concept of ‘real world’ applicability in authentic assessment literature and practice. The essay analyses student assessment, self-evaluation and survey responses and evaluates the implementation of authentic assessment in a humanities discipline through a framework structured by five key traits: critical thinking; interpersonal skills; engagement; ontological knowledge and social impact.
- PaperAssessment & Evaluation in Higher Education3 Jul 2026
AI-resistant and AI-resilient assessment in higher education: a systematic review of validity-grounded strategies, institutional frameworks, and equity implications
Asrat Genet Amnie
This systematic review differentiates between AI-resistant and AI-resilient assessment strategies in higher education. It synthesizes 47 peer-reviewed studies and 34 policy documents to propose a Six-Layer AI-Resilient Assessment Stack grounded in validity theory. The review finds that AI-detection tools are insufficient and highlights challenges in equity, reliability, and faculty development.
Original abstract
The emergence of large language models has precipitated a fundamental disruption to higher education assessment. Conventional instruments are susceptible to AI-assisted completion, threatening construct validity across disciplines and institutional contexts. When submitted work reflects AI capability rather than student competency, the inferential chain from performance to qualification is invalidated – a problem of design rather than detection. This systematic review makes two contributions. First, it theorises and grounds the distinction between AI-resistant assessment, which seeks to prevent or impede AI access through containment, and AI-resilient assessment, which assesses genuine human cognitive performance of intrinsic educational value regardless of the AI tools that exist. Second, it synthesises peer-reviewed empirical literature, policy documentation, and sector guidance published between 2022 and 2025. A PRISMA 2020-compliant search of five databases was conducted; two independent reviewers screened all records with substantial inter-rater agreement. Forty-seven peer-reviewed studies, 23 policy documents, and 11 sector reports were included. Six evidence-based assessment categories were identified, taxonomised, and evaluated against a structured adversarial threat taxonomy. AI-detection tools are structurally insufficient as primary responses. A Six-Layer AI-Resilient Assessment Stack, grounded in multi-trait multi-method validation logic, is proposed as an integrated institutional framework. Persistent challenges include equity, reliability, faculty development, and data protection.
- PaperarXiv — AI in Education (cs.CY)29 Jun 2026
Demystify, Use, Reflect, Assess (DURA): An Experience Report on LLM Integration in CS2
Margaret Ellis, Nikitha Donekal Chandrashekar, Sehrish Basir Nizamani, Mohammed Farghally et al.
The DURA framework (Demystify, Use, Reflect, Assess) guides the integration of Large Language Models into a CS2 course, including demystification, guided use with attribution, reflective activities, and adjusted assessments. Students reported positive engagement with LLMs while still valuing traditional support, and perceptions of instructor care improved.
Original abstract
Student access to Large Language Models (LLMs) is reshaping learning behaviors; at the same time students are entering the workforce where effective LLM use is becoming an expected skill. In this Experience Report we share our DURA framework (Demystify-Use-Reflect-Assess) and materials we used to restructure our CS2 course to allow the use of LLMs. We first demystified LLMs, then provided guidance on use with required attribution. We also added reflections related to LLM use at three points throughout the semester to encourage student meta-cognition around LLM use. We increased the value of proctored assessments in tandem with allowing retakes and including questions that explicitly assess skills from programming assignments. Students reported using LLMs for clarifying course concepts, debugging, understanding assignment guidelines, and determining test cases, but also still sought assistance via office hours and TAs, monitored Piazza, and reviewing course content. Students articulated thoughtful and strategic approaches to LLM use and also valued the instructional content and guidance from course staff. Student use of office hours increased slightly this semester and student perceptions that the instructor cares about them and their learning improved.
- PaperAssessment & Evaluation in Higher Education8 Jun 2026
Authentic assessment in higher education: conceptual evolution, key debates and implications for practice
Yusuf Josiah
This conceptual review traces the evolution of authentic assessment in higher education from performance-oriented formulations in the late 1980s to contemporary emphases on sustainability, digital mediation, and lifelong learning. It synthesizes historical and theoretical developments, highlights key debates around employability, equity, standardization, and academic integrity, and provides a foundation for future empirical research.
Original abstract
Authentic assessment has become increasingly central in higher education, reflecting a shift away from traditional, decontextualised testing toward assessment practices emphasising meaningful learning, integrated competence and the application of knowledge in context. Despite its growing prominence, authentic assessment remains conceptually fragmented, encompassing task-focused, competence-oriented, and more recent embedded and future-oriented interpretations. This article presents a comprehensive conceptual review tracing the evolution of authentic assessment from its early performance-oriented formulations in the late 1980s to contemporary conceptualisations emphasising sustainability, digital mediation, ethical engagement with emerging technologies and lifelong learning capabilities. The article synthesises historical and theoretical developments to examine how authenticity has been reframed in response to changing educational priorities and institutional contexts. It also highlights key debates in authentic assessment and tensions surrounding employability, equity, standardisation, digitally mediated learning environments and academic integrity. By providing an analytically structured account of the conceptual evolution of authentic assessment, the article clarifies conceptual ambiguities, situates contemporary developments within a broader historical trajectory and provides a foundation for future empirical research on the mechanisms, implementation and outcomes of authentic assessment across diverse higher education contexts.
- PaperAssessment & Evaluation in Higher Education6 Jun 2026
Assessment validity in the age of generative artificial intelligence: a critical review
Ramiz Ali, Jerry Maroulis
A critical review of literature on assessment validity in the age of generative AI identifies three dominant discourses: risk-oriented framings of academic integrity breaches, rule-based approaches focusing on detection and policy, and design-focused approaches advocating resilient assessment redesign.
Original abstract
Generative artificial intelligence (GenAI) presents significant challenges to assessment validity in higher education. In response, universities worldwide have invested heavily in risk-mitigation strategies. However, such efforts often prove insufficient given the rapid and on-going development of GenAI capabilities, prompting sustained debate about how assessment in higher education can remain valid in this evolving landscape. At the same time, inconsistencies in assessment design and implementation have raised concerns regarding the assurance of learning. This study presents a critical review of the literature on assessment validity in the context of GenAI through a discourse-analytic examination of articles published in 10 leading higher education journals between 2023 and June 2025. Articles were selected using predefined inclusion criteria, and discourse analysis was employed to synthesise dominant narratives within the literature. Three main discourses were identified: (1) risk-oriented framings that emphasise academic integrity breaches, (2) rule-based approaches centred on detection and policy enforcement, and (3) design-focused approaches that advocate resilient assessment redesign. The findings offer conceptual and practical implications for assessment practice and institutional policy in higher education in the age of GenAI.
- PaperAssessment & Evaluation in Higher Education4 Jun 2026
When ‘good teaching’ isn’t enough: learning environments that affect student feedback literacy
Caroline Xin Liu, Lily M. Zeng
A mixed-methods study of 547 Mainland Chinese undergraduates in Hong Kong found that assessment for understanding and clear goals significantly predicted feedback literacy, while good teaching and teacher feedback did not. Constructive alignment in programme-level learning environments, rather than teaching quality alone, is crucial for cultivating sustainable feedback literacy.
Original abstract
Despite the growing interest in student feedback literacy for student learning, how it is shaped within programme-level learning environments and its influence on learning outcomes remain underexplored. Even fewer studies have focused on how non-local students may experience this despite the key factors that were reported to have an impact on student feedback literacy would be different in their case. Using mixed methods, this project investigates the relationship among programme-level learning environments, feedback literacy, and learning outcomes. Study 1 analysed quantitative data from 547 Mainland Chinese undergraduates in Hong Kong. Structural equation modelling revealed that assessment for understanding and clear goals and standards significantly predicted feedback literacy, which mediated the effects of learning environments on learning outcomes, whereas good teaching and teacher feedback were not significantly associated with student feedback literacy. Study 2, based on interviews with fifteen students, indicated that aligned assessment designs and transparent standards were the key factors that enhanced student feedback literacy, explaining why teaching-related aspects may be less directly involved in creating affordances. The findings advance understanding of feedback literacy as influenced by the learning environments in the higher education context. The study highlights the critical role of constructive alignment in cultivating sustainable feedback literacy and improving student learning outcomes.