ai-in-assessment
Filtering by topic ai-in-assessment(5)Clear all filters
- PaperEdArXiv (OSF Preprints)20 Jul 2026
The Forcing Function: Ownership, Provenance, and Oral Defense in AI-Assisted Student Work
Greg O'Keefe
This paper presents an assessment redesign using oral defense to ensure student ownership in AI-assisted work. The framework assigns AI a narrow role (locating material) while requiring students to verify sources and construct an independent position, tested via a two-probe oral defense. The approach is designed for secondary classrooms with small class sizes and instructor familiarity.
Original abstract
Institutional responses to generative AI have largely relied on detection. Detection has failed, and Papers One and Two of this series argue that the failure is structural rather than technical: a finished piece of writing does not carry reliable evidence of the process that produced it, and no improvement in detection tools alters this. This paper proposes a different response. Rather than attempting to establish after the fact whether a student used AI, it redesigns the assessment so that success depends on a capacity AI cannot supply on the student's behalf. The instrument is oral defense, though not as a means of detecting deception. Oral defense is analyzed here as a forcing function: it alters what a student must do in advance in order to succeed, and therefore shapes behaviour before the assessment rather than adjudicating it afterward. The paper develops three connected components. The first is a research protocol that assigns AI a narrow role, the location of material, and reserves for the student the two operations that constitute ownership: verifying what each source argues, and constructing a position from the verified set. The second is a rule of exclusion: a claim whose primary source the student cannot obtain and read may not function as a premise in the argument, even in hedged form. The third is a two-probe defense that tests these two operations separately, incorporating a live friction point that cannot be anticipated and scripted. Ownership, on this account, is not authorship of the prose. It is verified provenance together with demonstrated independent synthesis. A student who used AI extensively and satisfies both conditions passes; a student who used no AI and satisfies neither fails. This is not a loophole in the framework but its central commitment. The paper is addressed to the secondary classroom rather than the university lecture hall, and the scope is integral to the argument rather than incidental. The mechanism developed here depends on an instructor who knows their students across a term, on class sections of roughly twenty to thirty, and on a developmental stage at which the knowledge being assessed is still forming rather than already established and merely being applied. Existing AI-resilient assessment frameworks, examined in Section 5, are designed for university cohorts in the hundreds, where that knowledge is assumed largely in place. The two scopes are not competing solutions to a single problem; they are solutions to different problems, and this paper claims only its own.
- PaperLanguage Testing19 Jul 2026
Can an AI Agent Replace Human Examiners in High-Stakes Interactive Speaking Tests? A Debate
Jing Xu, Lynda Taylor, Xiaoming Xi, Yasin Karatay et al.
This viewpoint article presents arguments for and against replacing human examiners with AI agents in high-stakes interactive speaking tests, drawing on a debate at the 2025 LTRC. It examines construct theory, practicality, washback, fairness, and ethical uses of AI in language assessment.
Original abstract
Generative artificial intelligence (GenAI) is advancing at a remarkable speed, and its promise in transforming current practice in language assessment has been articulated by applied linguistics researchers. An emerging application of GenAI is to integrate the technology into Spoken Dialogue Systems (SDSs) to simulate human interlocutors for the purpose of speaking practice or assessment. Despite rapid technological advances, the issue of whether an AI agent can replace a human examiner in one-on-one, high-stakes interactive speaking tests remains contentious. Building on a lively academic debate on this topic at the 2025 Language Testing Research Colloquium (LTRC) in Bangkok, this Viewpoint presents arguments both for and against this proposition in terms of construct theory, practicality, washback, fairness, ethical uses of AI, and so forth.
- PaperAssessment & Evaluation in Higher Education4 Jun 2026
University instructors’ contemporary assessment literacy: development and validation of a questionnaire
Goudarz Alibakhshi
The study developed and validated a 35-item questionnaire measuring university instructors' contemporary assessment literacy across nine dimensions, including learning-oriented assessment, feedback, learner involvement, ethics, digital and AI-responsive assessment, inclusivity, consequential validity, and learning analytics. The instrument was refined through expert review and administered to 462 Iranian instructors, with factor analysis supporting the multidimensional structure. This tool addresses gaps in existing measures by capturing modern assessment demands such as AI and learning analytics.
Original abstract
Assessment literacy has become a key professional competence in higher education, where instructors are expected to design learning-oriented, ethical, inclusive, digitally mediated and evidence-informed assessment practices. However, existing measures do not fully capture contemporary demands related to feedback, learner involvement, artificial intelligence, learning analytics, accessibility and assessment consequences. This study developed and validated a questionnaire measuring university instructors’ contemporary assessment literacy. Using a multiphase, mixed-methods instrument development design, the study was conducted in two sequential phases. In Phase 1, a preliminary 39-item pool was reviewed by 22 assessment experts from three universities in Tehran. Expert ratings supported item relevance, clarity, representativeness and essentiality, and 12 items were revised. In Phase 2, the revised questionnaire was administered to 670 university instructors from four Tehran universities; 462 usable responses were returned. Exploratory factor analysis supported a nine-factor solution explaining 66.66% of the variance. The retained dimensions were learning-oriented assessment literacy, feedback literacy, learner involvement, fairness and ethics, digital assessment literacy, AI-responsive assessment literacy, inclusive and accessible assessment literacy, consequential validity and washback literacy, and assessment data and learning analytics literacy. Confirmatory factor analysis supported the final 35-item model, with satisfactory reliability, convergent validity, discriminant validity and model fit.
- PaperAssessment & Evaluation in Higher Education29 May 2026
Between delegation and responsibility: an exploratory case study of graduate educators’ conceptualizations of AI-supported assessment using the AI assessment scale
Armağan Ateşkan
Graduate educators used the AI Assessment Scale (AIAS) as an ethical boundary-setting practice rather than a technical classification, negotiating the delegation of evaluative responsibility between human and AI agents. Their AI assessment literacy was stronger in procedural and experiential ethics than in structural ethics, and constructive alignment was most robust when AI was constitutively embedded in assessment design. Notably, designing without AI tools increased educators' awareness of habitual AI dependence and confidence in unassisted design.
Original abstract
This study investigates how graduate educators conceptualize and apply the AI Assessment Scale (AIAS) within their assessment design practice. Drawing on an exploratory qualitative case study design, the study analyzed eight AI-integrated lesson plans produced by in-service teachers in a technology elective course, supplemented by semi-structured interviews with four participants. Findings suggest that AIAS level selection is experienced not as a technical classification but as an ethical boundary-setting practice, a judgment about delegating evaluative responsibility between human and AI agents. Participants demonstrated variably developed AI assessment literacy: procedural ethics (integrity, authorship) and experiential ethics (learner agency) were more fully articulated than structural ethics (algorithmic bias, data governance). Constructive alignment was strongest when AI was constitutively embedded in the design; conversely, AI integration could shift assessment criteria from content-driven toward procedurally driven evaluation, a drift originating at the design-imagining stage before any tool was deployed. Notably, designing without AI tools heightened participants’ awareness of habitual AI dependence and, in several cases, increased confidence in unassisted design, suggesting that AI assessment literacy may require experiential constraint as well as conceptual instruction. Implications are discussed for teacher education, AIAS professional development, and AI assessment literacy.
- PaperETS Research Report Series14 May 2026
AutoSSD: A System for Automated Detection of Similar Speech Responses in Language Tests
Michael Fauss, Jiangang Hao, Chen Li, Michelle Palmer et al.
AutoSSD uses AI transcription to convert spoken responses to text and calculates verbatim similarity metrics to flag potential cheating via pre-scripted responses, providing an interface for expert review. The system proved effective in practice, offering insights for test security in spoken language assessments.
Original abstract
Evaluating spoken language proficiency stands as a pivotal component of language assessment. A well-known method for cheating in spoken language tests involves reciting pre-scripted responses. This paper describes a system developed to detect this type of cheating, called the Automated Speech Similarity Detector (AutoSSD). AutoSSD leverages artificial intelligence (AI) transcription to convert spoken responses into texts. It then calculates various metrics of verbatim similarity between the texts, flags pairs whose similarity levels exceed set thresholds, and provides an interactive interface for a subsequent expert review. Our system has proven effective in practical operation, and our findings can help other spoken language assessments bolster their test security measures. Suggested citation: Fauss, M., Hao, J., Li, C., Palmer, M., & Choi, I. (2026). AutoSSD: A system for automateddetection of similar speech responses in language tests (Research Memorandum No. RM-26-02). ETS. https://doi.org/10.64634/1g0whg02