llm-benchmark
Filtering by topic llm-benchmark(2)Clear all filters
- PaperarXiv — AI in Education (cs.CY)23 Jul 2026
AI Assistants Overassist
Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
Large language models acting as tutors often intervene too early and too frequently, providing complete solutions rather than targeted hints, according to a study introducing Int-Bench, a simulation-based benchmark for evaluating AI interventions during problem-solving. The research compared LLM teachers to humans across code debugging, math, and brain teasers, and found that LLMs prioritize short-term success over fostering deeper learning and generalization.
Original abstract
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assistants can scaffold thinking and foster learning, such benefits depend on how they help--for instance, intervening too early or too frequently may hinder true learning and cognitive engagement. Yet how AI systems navigate intervention decisions during problem-solving remains poorly understood. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions during learning. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains--code debugging, mathematics, and brain teasers--we evaluate LLM teachers on the frequency and timing of interventions, as well as their impact on both immediate task success and generalization to new problems. We also compare LLMs to humans, finding that LLMs intervene more frequently and earlier than humans. Moreover, in contrast to humans, they tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than supporting the reasoning processes needed for deeper learning and long-term success.
- PaperarXiv — Language & NLP (cs.CL)6 May 2026
Assessing Cognitive Effort in L2 Idiomatic Processing: An Eye-Tracking Dataset
Eduardo Santos, Juliana Carvalho, César Rennó-Costa
This paper presents an eye-tracking dataset to investigate how L2 learners process idiomatic expressions, revealing that L2 speakers often adopt a literal-first approach with measurable cognitive costs. The dataset, recorded from Portuguese L1 speakers of English across all CEFR levels using 60 Hz hardware, validates an inverse correlation between proficiency and regressive eye movements. It serves as a cognitively grounded benchmark for evaluating human processing models and alignment of large language models with human-like figurative understanding.
Original abstract
This paper presents the development and validation of an eye-tracking dataset designed to investigate how second-language (L2) learners process idiomatic expressions. While native speakers often rely on direct retrieval of figurative meanings, L2 speakers frequently adopt a literal-first approach, which incurs measurable cognitive costs. This resource captures these costs through ocular metrics recorded from Portuguese L1 speakers of English across all CEFR proficiency levels (A1-C2). Although the study uses entry-level 60 Hz hardware (Tobii Pro Spark), we demonstrate that this sampling rate provides sufficient data density to detect macro-cognitive events such as fixations and regressions in reading. Preliminary analysis validates the dataset by revealing a strong inverse correlation between language proficiency and regressive eye movements. Integrated into the MIA (Modeling Idiomaticity in Human and Artificial Language Processing) initiative, this dataset serves as a cognitively grounded benchmark for evaluating both human processing models and the alignment of large language models with human-like figurative understanding.