assessment-validity
Filtering by topic assessment-validity(3)Clear all filters
- PaperAssessment & Evaluation in Higher Education4 Jul 2026
Observer-based assessment in authentic management simulations: multi-source validity evidence and construct boundaries
Joerg Hruby
This study validates observer-based live assessments in authentic business simulations using Messick's unified validity framework. Multi-source evidence from 42 students in international strategy battles showed strong internal consistency and convergence with peer ratings, indicating that observer ratings effectively capture socially visible strategic performance rather than latent cognition.
Original abstract
Business schools increasingly use experiential simulations to develop applied strategic thinking, intercultural communication, and managerial decision-making. However, the validity of observer-based live assessments remains under-theorised, as existing research has focused mainly on written assignments and presentations. Following Messick’s unified validity framework, this study conceptualised observer ratings as score interpretations that require empirical justification from multiple sources of evidence to be valid. A multi-source, mixed-methods design integrated observer rubric evaluations, peer assessments, pre- and post-program surveys, and reflective reports from multicultural student teams participating in international business strategy battles involving headquarters–subsidiary negotiations (treatment: N = 42; control: N = 31; peer ratings: N = 412; reflection reports: N = 80). The results indicated strong internal consistency, meaningful convergence between observer and peer ratings, and substantial developmental gains across the program’s duration. Convergence patterns suggest that observer and peer systems capture the socially visible dimensions of enacted strategic performance more effectively than latent cognitive processes. The findings position observer-based assessments as a structured validity argument for externally enacted strategic performance in authentic higher-education simulations.
- PaperLanguage Testing25 Jun 2026
Reflections on the Practical Implementation of Knoch and Fan’s (2024) Good Practice Principles for Score Concordance Studies
Spiros Papageorgiou, Tony Clark
This viewpoint reflects on Knoch and Fan's (2024) good practice principles for score concordance studies, drawing on a large-scale study comparing IELTS Academic and TOEFL iBT scores. The authors emphasize methodological rigor, transparency, and construct comparability while noting limitations of concordance tables for admissions decisions, and recommend fair score requirements regardless of test chosen.
Original abstract
When different tests are used for the same purpose, score requirements should be comparable so that examinees cannot obtain an unfair advantage simply because of the test they chose. Drawing on our experience conducting a large-scale concordance study to allow for an empirical comparison of IELTS Academic and TOEFL iBT test scores, we review Knoch and Fan’s evaluative framework, explore methodological best practices and challenges, and offer future directions for score concordance research. We emphasize the importance of methodological rigor in collecting test-taker score data, transparency in analyzing such data to build score concordance tables, and a reasonable degree of construct comparability as a prerequisite for conducting a score concordance study, while also highlighting the limitations of concordance tables as standalone tools for admissions decisions. We note that some aspects of Knoch and Fan’s good practice principles are more straightforward to implement in practice than others. The good practice principles could be updated or adjusted after real-world application, which we describe with a view to furthering best practice in concordance research. We conclude this viewpoint with recommendations for decision-making that are based on fair score requirements, irrespective of which test the examinees chose.
- PaperAssessment & Evaluation in Higher Education6 Jun 2026
Assessment validity in the age of generative artificial intelligence: a critical review
Ramiz Ali, Jerry Maroulis
A critical review of 2023-2025 higher education journal articles identifies three dominant discourses on assessment validity in the age of generative AI: risk-oriented framings emphasizing academic integrity, rule-based approaches focused on detection and policy enforcement, and design-focused strategies advocating assessment redesign. The study finds current risk-mitigation efforts insufficient and offers conceptual and practical implications for maintaining assessment validity amid rapid GenAI developments.
Original abstract
Generative artificial intelligence (GenAI) presents significant challenges to assessment validity in higher education. In response, universities worldwide have invested heavily in risk-mitigation strategies. However, such efforts often prove insufficient given the rapid and on-going development of GenAI capabilities, prompting sustained debate about how assessment in higher education can remain valid in this evolving landscape. At the same time, inconsistencies in assessment design and implementation have raised concerns regarding the assurance of learning. This study presents a critical review of the literature on assessment validity in the context of GenAI through a discourse-analytic examination of articles published in 10 leading higher education journals between 2023 and June 2025. Articles were selected using predefined inclusion criteria, and discourse analysis was employed to synthesise dominant narratives within the literature. Three main discourses were identified: (1) risk-oriented framings that emphasise academic integrity breaches, (2) rule-based approaches centred on detection and policy enforcement, and (3) design-focused approaches that advocate resilient assessment redesign. The findings offer conceptual and practical implications for assessment practice and institutional policy in higher education in the age of GenAI.