interactional-competence
Filtering by topic interactional-competence(2)Clear all filters
- PaperLanguage Testing8 Jul 2026
“I Feel Like Talking With a Friend”: Exploring the Potential of Empathetic Spoken Dialogue Systems in Assessing Interactional Competence in L2 Oral Assessment
Lianzhen He, Ruixue Liang, Yang Zhao
The study compared an empathetic spoken dialogue system (E-SDS) with a neutral one (N-SDS) to assess interactional competence in L2 oral assessment. Results showed that the E-SDS elicited higher frequencies of interactional competence features and was perceived by learners as a competent, trustworthy, and emotionally supportive interlocutor, though technical limitations were noted.
Original abstract
As an important area of exploration within spoken dialogue systems (SDSs), empathetic spoken dialogue systems (E-SDSs) can provide emotional support for interlocutors, and this feature has the potential to be embedded in language teaching, learning, and assessment. However, the potential of E-SDSs in second language (L2) oral assessment remains underexplored, particularly regarding their ability to elicit interactional competence (IC) and learners’ perceptions of the two systems. To address these gaps, this study compares learners’ interaction with a neutral spoken dialogue system (N-SDS) with limited empathetic capability as a reference condition and an E-SDS, with an aim to explore its potential for L2 oral assessment. Twenty-five L2 learners completed two tasks (E-SDS and N-SDS). Their oral performances between the two tasks were examined in terms of IC features, and their perceptions of the E-SDS were also investigated through semi-structured interviews. Results indicated the E-SDS tends to elicit higher frequencies of certain IC features. Learners generally perceived the E-SDS as a competent, trustworthy, and emotionally supportive interlocutor, but certain concerns were also expressed concerning system design and technical limitations.
- PaperLanguage Testing8 Jul 2026
Assessing Interactional Competence Through Generative AI: Comparing Large Language Models as AI Interlocutors in the Paired Oral Discussion Test
Inyoung Na
This study compared GPT-4o and Claude 3.5 Sonnet as AI interlocutors in a paired oral discussion test to assess interactional competence (IC). Claude outperformed GPT-4o in eliciting IC features such as communication breakdown strategies and stance maintenance, and was perceived as more authentic by test takers. The findings underscore the importance of construct-driven evaluation when selecting large language models for language assessment contexts.
Original abstract
Interactional competence (IC) is essential for oral communication assessment, yet human partner variability can introduce construct-irrelevant variance in paired speaking tests. As an alternative to a test with a human interlocutor, this study describes the development of a large language model (LLM)-driven Spoken Dialogue System and compares GPT-4o to Claude 3.5 Sonnet to inform model selection for IC assessment. Twelve international students completed paired discussion tasks with both LLMs in counterbalanced order. System performance was evaluated through breakdown activation consistency, stance maintenance, and persona adherence. Test-taker performances were analyzed using interactional discourse analysis to identify IC features across three dimensions: topic management, interactional management, and interactive listening. Semi-structured interviews explored test takers’ perceptions of the AI partners. Results showed Claude outperformed GPT-4o in eliciting IC features, successfully activating communication breakdown strategies and maintaining oppositional stance, thereby creating more opportunities for test takers to demonstrate key IC abilities. Test takers perceived Claude as more authentic and natural, while GPT was perceived as more artificial. These findings demonstrate that different LLMs create distinct interactional conditions affecting both IC elicitation and test-taker perceptions. The findings highlight the need for construct-driven evaluation criteria when selecting LLMs for language-assessment contexts.