language-assessment
Filtering by topic language-assessment(22)Clear all filters
- PaperIELTS Partnership Research Reports
An investigation of the language assessment interests and needs of professional registration bodies in the UK: An unconsidered perspective?
This study explores the language assessment interests and needs of professional registration bodies in the UK, highlighting a previously overlooked perspective in the field.
- PaperIELTS Partnership Research Reports
Comparing New TOEFL 2026 with Former TOEFL 2023 and IELTS
This paper compares the upcoming TOEFL 2026 test with its 2023 predecessor and the IELTS, highlighting key differences in structure and scoring.
- PaperOpenAlex — TESOL researchForthcoming · 1 Dec 2026
Future of AI in English Language Education: Trends and Predictions
Prof. Jagadeesh Nerlekar, K Munianjinappa
This paper synthesizes current AI developments in English language education, identifying key trends such as personalized feedback, adaptive pathways, automated formative assessment, realistic speaking/listening practice via conversational agents, and AI-assisted material creation. It predicts that hybrid systems combining large language models with pedagogical scaffolding and teacher mediation will have the most near-term impact. The paper also addresses challenges including bias, privacy, over-reliance on automation, and the need for teacher training, offering recommendations for educators and policymakers.
Original abstract
Artificial intelligence (AI) is transforming English language education (ELE) by enabling personalized learning, automated assessment, adaptive content generation, and immersive practice environments. This paper synthesizes current developments, identifies emergent trends, and offers evidence-informed predictions about how AI will shape classroom practice, curriculum design, assessment, teacher roles, and policy over the next decade. Drawing on interdisciplinary literature from computer-assisted language learning (CALL), intelligent tutoring systems (ITS), natural language processing (NLP), and educational policy, the paper argues that the most significant near-term impact will stem from hybrid systems that combine large language models (LLMs) with pedagogically informed scaffolding and teacher mediation. Key trends discussed include (1) ubiquitous personalized feedback and adaptive pathways; (2) automated, formative assessment with rich analytics; (3) realistic speaking/listening practice via multimodal conversational agents and immersive virtual environments; (4) AI-assisted material creation and differentiation for diverse learner needs; and (5) data-driven teacher support and professional development. Predictions address likely improvements in scalability and access, as well as persistent challenges: bias and fairness in language models, privacy and data governance, over-reliance on automated feedback, and the need for robust teacher training and curricular alignment. The paper concludes with practical recommendations for educators, institutions, and policymakers to harness AI’s affordances while safeguarding equity, transparency, and pedagogical quality. These include adopting hybrid human–AI workflows, emphasizing explainability and interpretability in tools, developing clear data-ethics policies, investing in teacher capacity building, and prioritizing research-practice partnerships. The analysis aims to be actionable for practitioners and decision-makers planning for an AI-augmented future of English language learning. Keywords: artificial intelligence, English language education, adaptive learning, large language models, assessment, teacher role, ethics
- PaperApplied Linguistics24 Jul 2026
INTRODUCING SECOND LANGUAGE ASSESSMENT
Lin Shi, Lianzhen He
An introduction to second language assessment is provided, covering key concepts and approaches.
- PaperarXiv — AI in Education (cs.CY)9 Jul 2026
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight, Danielle Carvalho et al.
L2-Bench, an open-source benchmark of 1000+ task-response pairs, measures LLM capabilities in applying learning experience design principles for second language education and assessment. Validated by over 200 practitioners, it evaluates 12 competencies across 31 subcompetencies using a rubric-based methodology. Results show Claude Opus 4.7 achieves the highest overall score (85.5%), but performance drops substantially on harder tasks (69.9% to 73.4%).
Original abstract
Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity: 4.42/5.00, criteria adequacy: 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks (69.9% to 73.4%). L2-Bench provides education stakeholders better methods to make more informed decisions about real-world AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.
- PaperDOAJ — Language assessment1 Jul 2026
Multilevel auditory and cognitive processing in post-stroke aphasia: associations with language performance using ABR, LLR, and P300: a cross-sectional observational study
Agit Şimşek, Nihal Sümeyye Ulutaş, Feyza Deniz Saman
The study compared auditory evoked potentials (ABR, LLR, P300) between post-stroke aphasia patients and healthy controls, finding significantly prolonged latencies across subcortical, cortical, and cognitive levels in aphasia. These latency delays indicate generalized neural slowing rather than reduced recruitment, suggesting that electrophysiological measures may complement behavioral language assessments in aphasia evaluation.
Original abstract
ABSTRACT BACKGROUND: Aphasia following stroke is primarily characterized by language impairment; however, accumulating evidence suggests that deficits in auditory and cognitive processing may also contribute to impaired language function. Auditory evoked potentials provide objective markers of neural processing across subcortical, cortical, and cognitive levels and may help clarify the neurophysiological mechanisms underlying post-stroke aphasia. OBJECTIVES: To investigate multilevel auditory processing and its relationship with language performance in individuals with post-stroke aphasia using Auditory Brainstem Responses (ABR), Late Latency Responses (LLR), and P300 potentials. DESIGN AND SETTING: Cross-sectional observational study conducted at a university hospital in Türkiye. METHODS: Twenty-nine individuals with post-stroke aphasia and 33 age-matched healthy controls were included. Language performance was assessed using the Aphasia Language Assessment Test and a standardized naming task. Auditory processing was evaluated using ABR, LLR (P1, N1, P2, N2), and P300 potentials recorded according to standard electrophysiological protocols. Group comparisons were performed for latency and amplitude measures, and correlation analyses were conducted to examine associations between electrophysiological parameters and language performance. RESULTS: Compared with healthy controls, individuals with aphasia demonstrated significantly prolonged latencies in ABR waves and interpeak intervals, as well as in all LLR and P300 components (p < 0.05). No significant group differences were observed in amplitude measures. Naming performance was significantly lower in the aphasia group. Although correlations between language performance and electrophysiological measures did not reach statistical significance, moderate negative trends were observed between naming scores and N2 and P300 latencies. CONCLUSION: Post-stroke aphasia is associated with delayed auditory processing across subcortical, cortical, and cognitive levels, reflecting generalized slowing rather than reduced neural recruitment. Latency-based auditory evoked potential measures may complement behavioral language assessments and support a multilevel auditory–cognitive framework for aphasia evaluation and rehabilitation.
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of chatbot simulated spoken interaction compared to natural human speech. Using a corpus of ChatGPT output and the British National Corpus, they found that chatbot production more closely resembles written than spoken language, with higher lexical density and fewer spoken features like stance markers. The framework is intended for use in developing and validating AI-powered conversational agents for language assessment.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.
- PaperLanguage Testing19 Jun 2026
Book Review: Assessment of Plurilingual Competence and Plurilingual Learners in Educational Settings: Educative Issues and Empirical Approaches Melo-PfeiferS.OllivierC. (Eds.), Assessment of Plurilingual Competence and Plurilingual Learners in Educational Settings: Educative Issues and Empirical Approaches. Routledge, 2024. 272 pp. ISBN 9781032011097 (hbk) US$200.00 ISBN 9781032011103 (sbk) US$55.99 ISBN 9781003177197 (ebk) US$41.99
Gordon Blaine West
A review of a 2024 edited volume on assessing plurilingual competence in educational settings, covering both theoretical and empirical approaches.
- PaperETS Research Report Series16 Jun 2026
Mapping TOEIC® Link™ Scores to the Common European Framework of Reference for Languages
Kathryn Hille, Jonathan Schmidgall, Saerhim Oh, Renka Ohta
This study maps TOEIC® Link™ assessment scores to the Common European Framework of Reference for Languages using expert judgment. The Bookmark method was applied for multiple-choice listening and reading items, while the Expected Task Score method was used for constructed-response speaking and writing items. Panelists reported high agreement and satisfaction with the cut scores, with minor post-study adjustments for consistency across ETS assessments.
Original abstract
This study maps TOEIC® Link™ assessment scores to levels of the Common European Framework of Reference for Languages based on the judgments of a panel of experts. The Bookmark method was used to map the TOEIC Link assessments containing multiple-choice items (listening and reading), whereas the Expected Task Score method was used to map the assessments containing constructed-response items (speaking and writing). The panelists reported high levels of agreement, comfort with the panelist-recommended cut scores, and satisfaction with the standard setting process. A few small poststudy adjustments were made in the interest of congruence among ETS English proficiency assessments with common items or overlapping task types. Suggested citation: Hille, K., Schmidgall, J., Oh, S., & Ohta, R. (2026). Mapping TOEIC® Link™ scores to the Common European Framework of Reference for Languages (Research memorandum No. RM-26-04). ETS. https://doi.org/10.64634/4amdns26
- PaperAssessing Writing2 Jun 2026
Reassessing automated essay scoring with large language models: Evidence from API and GUI interfaces
Jieun Kim, Daniel Holden
This study re-examines the effectiveness of automated essay scoring using large language models by comparing results from API and GUI interfaces, providing evidence on the impact of interface design on scoring accuracy.
- PaperDOAJ — Language assessment1 Jun 2026
Glocalizing translation assessment: insights from the College English Test
Lin Zhang, Yan Jin
Using the College English Test (CET) in China as a case study, this article demonstrates how translation assessment can be glocalized – aligning test content with local socio-cultural contexts while maintaining global quality standards. The study analyzes the evolution of the CET translation component, showing how Chinese themes and discourse structures are integrated, and provides empirical validity evidence for the revised tasks.
Original abstract
Abstract This article examines the glocalization of translation assessment through a case study of the translation component of the College English Test (CET) in China. Specifically, the study investigates how the CET, through the reform of its translation component, sought to align its test content with the local context while maintaining global standards of quality. Our case study reveals that the content and format of the CET translation assessment have evolved in response to the changing societal and educational needs. Through an in-depth analysis of task design, the study illustrates how Chinese socio-cultural themes and discourse structures are incorporated to enhance the local relevance of the test. At the same time, the quality assurance measures implemented by the test developer are described in detail, and empirical evidence is presented to support the validity of the revised task. Our study illuminates how large-scale language assessments can integrate national education priorities with international standards of quality and fairness, thereby offering insights for language assessment professionals seeking to reconcile local responsiveness with global quality expectations.
- PaperETS Research Report Series29 May 2026
Alignment of the TOEIC Bridge® Tests With the Educational Functioning Level Descriptors for English as a Second Language
Mikyung Kim Wolf, Renka Ohta
This alignment study evaluated how well the TOEIC Bridge tests correspond to the NRS educational functioning level descriptors for ESL. Results showed strong content alignment across test forms, providing validity evidence for using these assessments in adult education programs for NRS accountability. The report offers methodology and recommendations for effective use.
Original abstract
This report presents the findings of an alignment study evaluating the extent to which the TOEIC Bridge® tests reflect the language knowledge, skills, and abilities described in the National Reporting System (NRS) educational functioning level descriptors for English as a second language (ESL). The study was intended to provide critical validity evidence supporting the use of these assessments in adult education programs serving ESL learners, particularly for purposes of NRS accountability. Results indicate strong evidence of content alignment across all test forms examined. The report details the research-based methodology and offers practical recommendations for the effective use of the TOEIC Bridge tests in the NRS. Suggested citation: Wolf, M. K., & Ohta, R. (2026). Alignment of the TOEIC Bridge® tests with the educational functioning level descriptors for English as a second language (Research Memorandum No. RM-26-03). ETS. https://doi.org/10.64634/1n934x31
- PaperETS Research Report Series20 Apr 2026
On the Representation of Racial and Ethnic Subgroups in AI-generated Texts: A Case Study in Automated Essay Scoring
Akshay Badola, Mo Zhang, Chen Li
This study examines whether LLMs can generate essays representing specific racial/ethnic subgroups and whether augmenting automated essay scoring (AES) training data with such generated texts reduces bias. Experiments with GPT-4 and GPT-4o show that while LLMs can produce subgroup-targeted essays, the inferred race/ethnicity distribution does not match the source data. Augmenting underrepresented groups with LLM-generated essays—regardless of whether the model correctly predicted race—improved human-score agreement and mitigated bias in AES systems.
Original abstract
In this study, we assess the capability of LLMs in generating essays of a specific race/ethnicity after being given example essays and rubric, and investigate the efficacy of data augmented in this manner for Automated Essay Scoring with respect to model performance and bias. In a series of experiments, we use models GPT-4 and GPT-4o, and ask them to generate essays from a given subgroup after inferring the race/ethnicity of the writer. We find that while LLMs can be directed to generate essays for specific demographic groups, the inferred racial and ethnic distribution in the generated data does not closely mirror the actual distribution observed in the source dataset. We augment existing data for underrepresented subgroups with LLM generated data separated into two groups with correct LLM race prediction and with incorrect race prediction and assess the improvement in agreement with human scores with quadratic weighted Kappa and bias mitigation as change in standardized mean difference. Our analysis shows that while LLMs struggle to predict the race accurately from given samples, augmentation with such data can be helpful to mitigate bias regardless. Suggested citation: Badola, A., Zhang, M., & Li, Chen. (in press). On the representation of racial and ethnic subgroups in AI-generated texts: A case study in automated essay scoring. ETS Research Report Series. https://doi.org/10.64634/ac01td58
- PaperLanguage Testing20 Apr 2026
A Systematic Review of Test Component Ordering with Implications for Language Assessment
Ben Naismith, Ramsey L. Cardwell
A systematic review of 88 studies from 1933 to 2023 examined how the ordering of test components (items, tasks, sections) affects language assessment outcomes. Findings suggest easy-to-hard ordering may improve performance, but effects are mediated by test-taker characteristics like anxiety and proficiency, while common practices like ordering by skill lack strong empirical support.
Original abstract
How to optimize ordering the parts of a language test (e.g., items, tasks, and sections) is often overlooked or assumed to be self-evident. However, ordering choices are an essential component of test design as they may impact test takers’ objective performance, perceived performance, or affective states. In this paper, we report on a systematic review of test component ordering research, focusing on 88 studies from 1933 to 2023. We provide a narrative synthesis, describing typical outcome variables (e.g., test-taker performance), independent variables (e.g., order of difficulty), and mediating variables (e.g., language proficiency). Key findings, mostly from higher educational contexts, indicate that easy-to-hard ordering may lead to better performance, though these effects are mediated by various test-taker characteristics, especially anxiety and proficiency. While section ordering by language skills is common practice in language testing, there is scant empirical support for this approach, or for organizing tests by content or format. We discuss the implications of these findings for language test developers and suggest avenues for future research, particularly the need for more studies on ordering effects in language assessment contexts.
- PaperETS Research Report Series4 Apr 2026
TOEIC® Link™ Assessments Technical Manual
Jaime Cid, Jonathan Schmidgall, Elizabeth Park
This technical manual provides a comprehensive overview of the TOEIC Link assessments, detailing their purpose, design, constructs, and tasks. It covers development processes for listening, reading, speaking, and writing components, emphasizing validity and reliability. The manual is intended as a living document for test users and stakeholders.
Original abstract
This technical manual provides a comprehensive overview of the TOEIC® Link™ assessments, offering detailed insights into their purpose, design, and intended users. The manual begins with an introduction that outlines the assessment objectives and target audience. Subsequent sections delve into the specific constructs and tasks of the four TOEIC Link assessments: Listening, Reading, Speaking, and Writing. The remaining sections examine the design and development processes for both the listening and reading, as well as the speaking and writing assessments, highlighting methodologies used to ensure validity and reliability. Together, these sections present a thorough guide for test users and stakeholders seeking to understand the TOEIC Link assessments and their application in measuring English language proficiency. Designed as a living document, this manual will be updated as the test’s design, administration, scoring, and evidence of measurement quality (including reliability, validity, and fairness) evolve, along with its intended uses. Suggested citation: Cid, J., Schmidgall, J., & Park, E. (2026). TOEIC® Link™ assessments technical manual (Research Report No. RR-26-04). ETS. https://doi.org/10.64634/7e2pzg04
- PaperETS Research Report Series31 Dec 2025
Using Ordinal Rescore Measures to Monitor Rater Drift
John Donoghue, Adrienne Sgammato
This study evaluates methods for monitoring rater drift when scoring constructed response items across occasions. It shows that an alternative analysis conditioning on the rescore design effectively detects drift, while the usual trend analysis produces biased estimates. Omnibus measures based on t-tests showed marginally higher power than those based on d-statistics.
Original abstract
When constructed response items are used on more than one occasion, a natural concern is whether the scoring is consistent (e.g., not more lenient or strict) across the occasions. It is common to conduct trend scoring, in which a set of Occasion A responses are rescored at Occasion B. The responses are usually selected according to some rescore design, such as being balanced (with an equal number from each score category), proportional to the distribution of Occasion A scores, or a mixed version of these two designs. Recent work has demonstrated that treating the two-way table as if it arose from multinomial sampling is incorrect and can yield seriously biased estimates of whether the scores are lower or higher at Occasion B. The present study builds on these results by incorporating ordinal measures of change. It contrasts the usual trend analysis with an alternative analysis that explicitly conditions on the rescore design and finds only the latter to be effective. Omnibus measures based on combining the individual t-tests or d-statistics are examined. Measures were somewhat conservative in Type I error control and had good power to detect drift. Omnibus measures based on t-tests had marginally higher power, having higher correct detection rates than those based on the d-statistic in 1%–8% of the cases. The difference between the best versions (E weighted, which is based on t-tests, vs. D weighted, which is based on d-statistics) was only 1.8%. Suggested citation: Donoghue, J. R., & Sgammato, A. (2025). Using ordinal rescore measures to monitor rater drift(Research Report No. RR-25-15). ETS.
- PaperDOAJ — Language assessment1 Dec 2025
Emerging successes and persistent challenges in Hungarian minority education in Romania
Imre Tódor
An analysis of 2020–2025 assessment data shows that a 1.3–1.5 point gap in Romanian language performance between Hungarian minority and Romanian majority students persists at the 8th-grade level, while math scores are nearly equal, confirming the gap is linguistic. Following the introduction of a differentiated 'Romanian as a non-native language' curriculum, 2025 baccalaureate pass rates improved in high-minority regions, especially in vocational schools, though effectiveness varies by region and school type.
Original abstract
This study examines the early impacts of recent curriculum and examination reforms in Romanian minority education, focusing on the introduction of the “Romanian as a non-native language” curriculum for Hungarian-speaking students. Using aggregated national assessment and baccalaureate data from 2020–2025, the research analyzes trends in Romanian language performance among minority students, compares results across regions and school types, and uses mathematics performance as a comparative indicator to contextualize language-specific achievement patterns. Descriptive, cohort-comparative, and proportion-difference analyses, complemented by hypothetical “what-if” calculations, reveal that while a persistent 1.3–1.5 point gap remains between minority and majority students in Romanian language performance at the 8th-grade level, mathematics scores are nearly equivalent, indicating that the gap is linguistic rather than cognitive. In the 2025 baccalaureate – the first year of full curriculum implementation – pass rates improved notably in high-minority regions (e.g., Harghita +5.3 pp, Covasna +1.6 pp), alongside a significant reduction in failure rates, particularly in vocational and technical schools. The findings suggest that aligning examination content with a differentiated curriculum may be associated with more favorable educational outcomes among minority students, though effectiveness varies by region and school type. Sustainable gains require targeted teacher training, adequate resources, and systematic monitoring to address persistent structural and contextual disparities.
- PaperETS Research Report Series19 Nov 2025
An Evaluation of Item Fit Based on Generalized Residual Item Response Functions
Xiangyi Liao, Peter Van Rijn, Sandip Sinharay
Developed a new method to evaluate item fit in item response theory (IRT) models by summarizing generalized residuals into a single statistic per item, accounting for estimation error. Simulations showed similar Type I error control to an existing statistic with slight improvements for small samples, but low power except for extreme misfit. The method combines prior work by Haberman and Kondratek.
Original abstract
Evaluation of item ft for item response theory (IRT) models often involves a comparison of the observed and expected item response functions (IRFs). Several statistics have been suggested for evaluating item ft based on the discrepancy between IRFs, but the asymptotic distributions of the statistics under the null hypothesis are often not well established. Haberman et al. developed a method for evaluating the ft of IRFs based on generalized residuals. These residuals are functions of the latent proficiency variable in the IRT model and follow the standard normal distribution asymptotically. We develop a method to summarize these generalized residuals into a single summary statistic for each item and evaluate its asymptotic distribution. Kondratek suggested a similar Wald-type statistic, but without accounting for the uncertainty in the estimation of the item parameters. Our method combines the work of Haberman and Kondratek, resulting in a single ft statistic per item while accounting for estimation error. A series of simulations was carried out to investigate the performance of our statistic and compare it to several popular item ft statistics. Our method resulted in similar Type I errors as Kondratek’s statistic, with slightly better results in the case of small samples. Furthermore, the recovery was consistent across different levels of item difficulty, and power of the new item ft statistic was relatively low, except for problematic individual items, but this result was found with two competing statistics as well. Suggested citation: Liao, X., van Rijn, P., & Sinharay, S. (2025). An evaluation of item fit based on generalized residual item response functions (Research Report No. RR-25-13). ETS. https://doi.org/10.64634/b68vz316
- PaperERIC — Assessment & second language1 Jan 2025
A Systematic Review of Differential Item Functioning in Second Language Assessment
Xueliang Chen, Vahid Aryadoust, Wenxin Zhang
This systematic review of 83 articles found that differential item functioning (DIF) analysis in second language (L2) assessment primarily relies on classical methods like Rasch, Mantel-Haenszel, and IRT, with emerging approaches such as cognitive diagnostic models also appearing. Most studies focused on manifest grouping variables (e.g., gender, language background) and receptive skills, and often lacked empirical justification for DIF causes, highlighting the need for improved practices and broader consideration of test-taker diversity.
Original abstract
The growing diversity among test takers in second or foreign language (L2) assessments makes the importance of fairness front and center. This systematic review aimed to examine how fairness in L2 assessments was evaluated through differential item functioning (DIF) analysis. A total of 83 articles from 27 journals were included in a systematic review. The findings suggested that classical DIF techniques were dominant in use, particularly Rasch-based methods, the Mantel-Haenszel procedure, item response theory (IRT) approaches, logistic regression, and SIBTEST, but emerging methods such as DIF analysis based on cognitive diagnostic models were also identified. Most DIF studies examined manifest grouping variables such as gender and language background and were based on assessments of receptive language skills such as reading and listening comprehension. DIF analyses were mostly conducted in an exploratory fashion and causes of DIF were often justified on speculative rather than empirical grounds. In addition, the quality of DIF analyses was undermined by suboptimal reporting practices. Our results suggest the need to improve current DIF practices, to consider alternative DIF detection methods aligning with emerging views of measurement bias, and to adequately account for the heterogeneity of L2 test takers. The findings have implications for test design and use, fairness, and validity in L2 assessments.
- PaperERIC — Assessment & second language1 Jan 2025
The Use of E-Portfolio as Innovative Creation of EFL Students in Language Assessment
Eny Syatriana, Saiful, Ariana, Adni Dzilarsy
Developed an e-portfolio measurement instrument model and surveyed 40 EFL students at an Indonesian university to evaluate its implementation. Findings revealed obstacles such as inadequate internet access and difficulty adjusting to the e-portfolio, while also outlining advantages and drawbacks for academic assessment. The model aims to improve learning outcomes and guide practitioners in higher education.
Original abstract
The main aim of this paper is to have an innovative model for helping students to design an e-portfolio assessment subject. This study developed an e-portfolio measurement instrument model, at the individual level of analysis, using responses from 40 e-portfolio student users from the fourth-semester students of Muhammadiyah University, were purposively selected to assess the success of their e-portfolio implementations from their students' perspective. Academic institutions can use the results of this research to assess the success of their e-portfolio implementations from their students' perspective. It employs the research and development method to collect data, which were used to achieve strengthen confidence in the learning outcomes. Innovative learning is an approach or learning method that uses new, creative ways, prioritizing skills visible in student learning outcomes. Language Assessment and Evaluation courses play an important role in various activities. In terms of the use of e-portfolio for teaching objectives to be achieved, lecturers are required to use interesting methods of e-portfolio implementation teaching assessment and evaluation skills because the results can be used to improve learning outcomes. The findings show that the students' common perceptions of obstacles were inadequate Internet access and trouble adjusting to the e-portfolio. The implementation of the e-portfolio is new, so it is vital to get the opinions of those who use it daily. The paper outlines the advantages and drawbacks of utilizing the e-portfolio as a tool for academic assessment. This will direct scholars and practitioners in the proper implementation of the e-portfolio in higher education establishments.
- PaperDOAJ — Language assessment1 Apr 2018
Die Versprachlichung räumlicher Relationen als Herausforderung im Erwerb des Deutschen als Erst- und Zweitsprache: Warum „Kommissar Wuschel“ in einer Sprachstands-App über Stämme steigt und auf Stühlen steht
Maike Schug
A language assessment app for German-speaking children aged 4;6 to 6 uses a serious game to test verbalization of spatial relations. The paper argues that tasks requiring unambiguous descriptions of object positions and movements effectively differentiate language competence levels, as supported by pre-test results.
Original abstract
Im Basismodul eines neuen, am natürlichen Kommunikationsverhalten von Vorschulkindern orientierten Verfahrens zur Sprachstandsdiagnose, der „Kommissar Wuschel“-App, kann im Rahmen eines Serious Game die Kompetenz von 4;6 bis 6 Jahre alten Kindern untersucht werden, räumliche Relationen zu versprachlichen. Der Beitrag begründet die Wahl dieser Art von Sprachhandlungsaufgaben unter Rekurs auf die Grundlagenforschung zu Erst- und Zweitspracherwerb, aber auch auf die überzeugenden Ergebnisse der Pilotierung. Die Pre-Tests haben gezeigt, dass die Herausforderung, Positionen von Gegenständen und Lebewesen im Raum sowie deren Bewegungen unmissverständlich sprachlich auszudrücken, ein gutes Mittel dafür ist, Kompetenzniveaus differenziert zu unterscheiden.A recently developed language assessment app, called „Kommissar-Wuschel“, uses a Serious Game Scenario to test aspects of language competence of 4;6 to 6;0 year old children. The main part of the app is particularly designed to investigate ways of expressing spatial relations and builds on basic research in first and second language acquisition. The paper argues that speech act tasks of this kind—aiming at describing how objects are located and move in natural environments—allow to differentiate levels of language competence on a fine-grained scale.
- PaperDOAJ — Language assessment1 Apr 2013
English language assessment in bilingual CLIL instruction at the primary level in Finland: Search for updated and valid assessment methods
ZIF Redaktion
This paper examines English language assessment practices in bilingual CLIL instruction in Finnish primary schools, with a focus on identifying updated and valid assessment methods. It addresses the need for assessment approaches that align with the integrated content and language learning goals of CLIL.