validity
Filtering by topic validity(4)Clear all filters
- PaperAssessment & Evaluation in Higher Education3 Jul 2026
AI-resistant and AI-resilient assessment in higher education: a systematic review of validity-grounded strategies, institutional frameworks, and equity implications
Asrat Genet Amnie
This systematic review distinguishes AI-resistant (preventing AI access) from AI-resilient (assessing genuine human cognitive performance) assessment in higher education. It proposes a Six-Layer AI-Resilient Assessment Stack grounded in validity logic and identifies six evidence-based assessment categories. The review finds AI-detection tools insufficient as primary responses and highlights challenges including equity, faculty development, and data protection.
Original abstract
The emergence of large language models has precipitated a fundamental disruption to higher education assessment. Conventional instruments are susceptible to AI-assisted completion, threatening construct validity across disciplines and institutional contexts. When submitted work reflects AI capability rather than student competency, the inferential chain from performance to qualification is invalidated – a problem of design rather than detection. This systematic review makes two contributions. First, it theorises and grounds the distinction between AI-resistant assessment, which seeks to prevent or impede AI access through containment, and AI-resilient assessment, which assesses genuine human cognitive performance of intrinsic educational value regardless of the AI tools that exist. Second, it synthesises peer-reviewed empirical literature, policy documentation, and sector guidance published between 2022 and 2025. A PRISMA 2020-compliant search of five databases was conducted; two independent reviewers screened all records with substantial inter-rater agreement. Forty-seven peer-reviewed studies, 23 policy documents, and 11 sector reports were included. Six evidence-based assessment categories were identified, taxonomised, and evaluated against a structured adversarial threat taxonomy. AI-detection tools are structurally insufficient as primary responses. A Six-Layer AI-Resilient Assessment Stack, grounded in multi-trait multi-method validation logic, is proposed as an integrated institutional framework. Persistent challenges include equity, reliability, faculty development, and data protection.
- PaperLanguage Testing26 Apr 2026
Justifying the Score or Informing the Stakeholder? Transparency Challenges in Large-Scale Language Testing
Vahid Aryadoust
Large-scale language testing operates as both an evaluative practice and a commercial enterprise, creating a tension between transparency and persuasive argument-based validity. This viewpoint proposes that test developers publish form-specific validity reports and publicly disclose test limitations, drawing on pharmaceutical industry norms to protect consumers in high-stakes contexts such as immigration and university admissions. The goal is to shift from justifying scores toward more open science practices.
Original abstract
Language testing is both an evaluative practice and a commercial enterprise shaped by market forces. Within this context, test developers have a responsibility to ensure transparency with test users, particularly when scores inform high-stakes decisions. This Viewpoint contrasts transparency in communicating the truth about language tests with the persuasiveness of argument-based validity, noting that persuasiveness, although central to such arguments, is not equivalent to transparency. Two measures are proposed to strengthen transparency. First, test developers should publish test-form-specific validity reports detailing content, psychometric properties, and interpretation guidelines for each form. Second, they should clearly explain a test’s limitations to the public, especially when scores are used in high-stakes settings, such as immigration or university admission, without adequate validation. The latter measure draws on regulatory norms in the pharmaceutical industry, where transparency can protect consumers from potential misuse. Specific steps are outlined to support these measures. Overall, these proposals aim to shift the emphasis from justification and persuasion toward transparency and align language testing practices more closely with open science principles.
- PaperETS Research Report Series4 Apr 2026
TOEIC® Link™ Assessments Technical Manual
Jaime Cid, Jonathan Schmidgall, Elizabeth Park
This technical manual from ETS describes the design, development, and validation of the TOEIC Link assessments for listening, reading, speaking, and writing. It outlines the constructs, tasks, and methodologies used to ensure reliability and validity, with updates planned as the test evolves. The manual serves as a guide for test users and stakeholders.
Original abstract
This technical manual provides a comprehensive overview of the TOEIC® Link™ assessments, offering detailed insights into their purpose, design, and intended users. The manual begins with an introduction that outlines the assessment objectives and target audience. Subsequent sections delve into the specific constructs and tasks of the four TOEIC Link assessments: Listening, Reading, Speaking, and Writing. The remaining sections examine the design and development processes for both the listening and reading, as well as the speaking and writing assessments, highlighting methodologies used to ensure validity and reliability. Together, these sections present a thorough guide for test users and stakeholders seeking to understand the TOEIC Link assessments and their application in measuring English language proficiency. Designed as a living document, this manual will be updated as the test’s design, administration, scoring, and evidence of measurement quality (including reliability, validity, and fairness) evolve, along with its intended uses. Suggested citation: Cid, J., Schmidgall, J., & Park, E. (2026). TOEIC® Link™ assessments technical manual (Research Report No. RR-26-04). ETS. https://doi.org/10.64634/7e2pzg04
- PaperETS Research Report Series9 Oct 2025
TOEFL iBT® Technical Manual
Venessa Manna, Shuhong Li, Spiros Papageorgiou, Lixiong Gu
The TOEFL iBT Technical Manual describes the test's purpose, design, and scoring procedures, and outlines a research agenda to support score interpretation and use. Updated for changes starting January 2026, it serves as a living document that will evolve with future evidence of reliability, validity, and fairness.
Original abstract
This technical manual describes the purpose and intended uses of the TOEFL iBT test, its target test-taker population, and relevant language use domains. The test design and scoring procedures are presented first, followed by a research agenda intended to support the interpretation and use of test scores. Given the updates to the test starting January 2026, this technical manual is intended to serve as an overview and rationale for the test design as well as a reference point for informing investigations of validity evidence to support the intended test uses over time. Designed as a living document, this manual will be updated as the test's design, administration, scoring, and evidence of measurement quality (including reliability, validity, and fairness) evolve, along with its intended uses. Suggested citation: Manna, V. F., Li, S., Papageorgiou, S., & Gu, L. (2025). TOEFL iBT® technical manual (TOEFL Research Report No. RR-106). ETS. https://doi.org/10.64634/eje8f497