corpus-linguistics
Filtering by topic corpus-linguistics(2)Clear all filters
- PaperLanguage Testing28 Jun 2026
Evaluating Chatbot Authenticity in Simulations of Spoken Interaction: Demonstrating The Utility of Corpus-Based Methods for Development and Validation
Dana Gablasova, Luke Harding, Vaclav Brezina, Emil T. Hazelhurst et al.
This study develops a corpus-based framework to evaluate the authenticity of chatbot-generated language compared to natural spoken interaction. Using a Chatbot Corpus (290,000 words) from ChatGPT (3.5 and 4) and the British National Corpus 2014, analyses at macro, meso, and micro levels found that chatbot output more closely resembles written than spoken genres, with higher lexical density and fewer spoken features like stance and pragmatic markers. The framework is intended for use in developing and validating conversational agents for language assessment and learning.
Original abstract
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.
- PaperReCALL6 May 2026
Tracing the diachronic effects of data-driven learning on lexical complexity in EFL learners’ argumentative writing
Yanan Zhao, Jihua Dong
Data-driven learning (DDL) instruction significantly improved lexical complexity in Chinese EFL learners' argumentative writing over five time points, while a non-DDL control group declined. Learners showed nonlinear individual trajectories in lexical complexity development and reported positive attitudes toward DDL, though some challenges in corpus use remained.
Original abstract
This study investigates the effectiveness of data-driven learning (DDL) in promoting lexical complexity in Chinese English as a foreign language (EFL) learners’ argumentative writing, tracks developmental trajectories, and examines learners’ perceptions. Adopting a quasi-experimental design, one class ( n = 26) received DDL instruction, and the other ( n = 22) received non-DDL instruction. Data were collected using triangulation, including argumentative writing samples from five time points, pre- and post-instruction questionnaires and semi-structured interviews. Results showed that learners in the DDL class significantly improved their lexical complexity, while the non-DDL class experienced declines. Across the five time points, nonlinear trajectories were observed in lexical complexity at the individual learner level. Learners reported positive attitudes toward DDL, though some challenges in corpus use remained. These findings provide empirical support for the effectiveness of DDL in promoting lexical complexity development in Chinese EFL learners’ argumentative writing and provide pedagogical implications for corpus-based writing instruction.