validity
Filtering by topic validity(4)Clear all filters
- PaperAssessment & Evaluation in Higher Education3 Jul 2026
AI-resistant and AI-resilient assessment in higher education: a systematic review of validity-grounded strategies, institutional frameworks, and equity implications
Asrat Genet Amnie
This systematic review differentiates between AI-resistant and AI-resilient assessment strategies in higher education. It synthesizes 47 peer-reviewed studies and 34 policy documents to propose a Six-Layer AI-Resilient Assessment Stack grounded in validity theory. The review finds that AI-detection tools are insufficient and highlights challenges in equity, reliability, and faculty development.
Original abstract
The emergence of large language models has precipitated a fundamental disruption to higher education assessment. Conventional instruments are susceptible to AI-assisted completion, threatening construct validity across disciplines and institutional contexts. When submitted work reflects AI capability rather than student competency, the inferential chain from performance to qualification is invalidated – a problem of design rather than detection. This systematic review makes two contributions. First, it theorises and grounds the distinction between AI-resistant assessment, which seeks to prevent or impede AI access through containment, and AI-resilient assessment, which assesses genuine human cognitive performance of intrinsic educational value regardless of the AI tools that exist. Second, it synthesises peer-reviewed empirical literature, policy documentation, and sector guidance published between 2022 and 2025. A PRISMA 2020-compliant search of five databases was conducted; two independent reviewers screened all records with substantial inter-rater agreement. Forty-seven peer-reviewed studies, 23 policy documents, and 11 sector reports were included. Six evidence-based assessment categories were identified, taxonomised, and evaluated against a structured adversarial threat taxonomy. AI-detection tools are structurally insufficient as primary responses. A Six-Layer AI-Resilient Assessment Stack, grounded in multi-trait multi-method validation logic, is proposed as an integrated institutional framework. Persistent challenges include equity, reliability, faculty development, and data protection.
- PaperLanguage Testing2 Jul 2026
Test Review: The Test of Proficiency in Korean
So-Young Lim, John Dylan Burton
This test review provides an overview of the TOPIK's history, purposes, design, and administration, and appraises its strengths and challenges. The lack of publicly available information on test construct, psychometric properties, and standard-setting procedures makes it difficult to fully evaluate score validity and reliability. Additional validation research is needed to examine score generalizability beyond academic domains.
Original abstract
As a nationally accredited test, the Test of Proficiency in Korean (TOPIK) assesses general Korean proficiency as a foreign language for various high-stakes purposes, such as university admissions, employment, and visa issuance. For this reason, the social impact of TOPIK on test takers is significant and cannot be underestimated. Despite the growing number of international test takers and stakeholders, there is limited validation research evaluating TOPIK’s various uses, as well as critical evaluation of the test itself. Thus, this test review provides an overview of the history, test purposes and use, design, and administration of TOPIK and offers an appraisal of its strengths and challenges. While the test scores are broadly utilized for their intended purposes and provide some evidence of language development across four skills in Korean, the lack of publicly available information on the test construct, psychometric properties, and standard-setting procedures makes it difficult to fully evaluate the validity and reliability of the scores. Furthermore, additional evidence and validation research are needed to examine the generalizability of TOPIK scores beyond academic domains, given the test’s diverse applications. Addressing these gaps is critical to meet the needs of various stakeholders and to strengthen the overall validity of the test.
- PaperDOAJ — Language assessment1 Jul 2026
The critical extension framework: making organisational assumptions explicit in multilingual writing assessment
Projnya Mojumdar
The study introduces the Critical Extension Framework (CEF) to address how ESL writing rubrics conflate rhetorical preference with language proficiency. Through critical discourse analysis of TOEFL, IELTS, CEFR, and classroom rubrics, it reveals that descriptor language encodes linear, thesis-first norms. The framework separates pattern recognition, rhetorical effectiveness, and language resources to improve score precision and fairness, with implications for multilingual writing assessment in Asian EFL/ESL contexts.
Original abstract
Abstract At the construct-articulation stage, current ESL writing rubrics quietly pull rhetorical preference into the operational definition of language proficiency. The interpretive consequences reach beyond measurement alone: once scores cannot tell organisational strategy apart from language control, teacher interpretation, feedback, and gatekeeping decisions all lose precision. This framework-development study traces the problem through critical discourse analysis of four major frameworks (TOEFL iBT, IELTS Academic Writing, CEFR, and classroom assessment principles). It shows how descriptor language governing organisation, coherence, and development encodes linear, thesis-first norms. Coding-scheme stability is established through intra-analyst (Cohen’s κ = 1.00) and inter-analyst (81.7 per cent agreement) checks on stratified subsets matched in size and scheme (12 of 34; 35.3 per cent). The Critical Extension Framework (CEF) names this diagnosis explicitly as pattern recognition folded into rhetorical effectiveness at the descriptor level. It responds by inserting a pattern-recognition step before scoring, separating the three inferential layers that current rubrics conflate: pattern recognition, rhetorical effectiveness, and Language Resources. Situated within established validity architecture at the domain-description inference, the framework makes organisational assumptions explicit before they enter later scoring, generalisation, and use claims. A preliminary two-rater pilot suggests that Language Resources scores may converge when raters work from the scoring descriptors and scoring protocol in Appendix B together with the essay texts. Rhetorical Effectiveness ratings on the reader-responsible implicit essay show pronounced, construct-sensitive disagreement. That disagreement locates the framework’s heaviest validation burden. The point holds with particular force for multilingual writing assessment in Asian EFL/ESL contexts.
- PaperLanguage Testing26 Apr 2026
Justifying the Score or Informing the Stakeholder? Transparency Challenges in Large-Scale Language Testing
Vahid Aryadoust
Large-scale language testing needs greater transparency, according to this viewpoint, which contrasts transparency with persuasiveness in argument-based validity. Two measures are proposed: publishing test-form-specific validity reports and clearly explaining test limitations, inspired by pharmaceutical industry norms. The goal is to shift emphasis from justification to transparency and align with open science.
Original abstract
Language testing is both an evaluative practice and a commercial enterprise shaped by market forces. Within this context, test developers have a responsibility to ensure transparency with test users, particularly when scores inform high-stakes decisions. This Viewpoint contrasts transparency in communicating the truth about language tests with the persuasiveness of argument-based validity, noting that persuasiveness, although central to such arguments, is not equivalent to transparency. Two measures are proposed to strengthen transparency. First, test developers should publish test-form-specific validity reports detailing content, psychometric properties, and interpretation guidelines for each form. Second, they should clearly explain a test’s limitations to the public, especially when scores are used in high-stakes settings, such as immigration or university admission, without adequate validation. The latter measure draws on regulatory norms in the pharmaceutical industry, where transparency can protect consumers from potential misuse. Specific steps are outlined to support these measures. Overall, these proposals aim to shift the emphasis from justification and persuasion toward transparency and align language testing practices more closely with open science principles.