multi-agent-systems
Filtering by topic multi-agent-systems(1)Clear all filters
- PaperComputers and Education: Artificial Intelligence22 Jun 2026
ASTRA: A synthetic benchmark for trace-based evaluation of socially intelligent multi-agent tutoring and participation-balanced collaboration in introductory programming
Solomon Sunday Oyelere
This paper introduces ASTRA, a benchmark framework for studying collaborative programming with socially differentiated agents, releasing a synthetic dataset simulating three tutoring configurations to evaluate interaction dynamics and participation balance. The findings are simulated evidence intended for benchmarking and reproducible pipeline development, not causal estimates.
Original abstract
Generative AI is rapidly entering introductory programming, yet evidence about how learners coordinate with AI, especially in dyads, remains limited, and open datasets that support reproducible, trace-based evaluation are scarce. I present ASTRA (Adaptive Socially-intelligent Team Reasoning Agents), a multi-agent tutoring prototype and benchmark framework for studying collaborative programming with socially differentiated agents. ASTRA supports three configurations: alone_tutor (one learner with a Tutor agent), pair_tutor (two learners with a Tutor agent), and pair_multiagent (two learners with Tutor and Facilitator agents, where the Facilitator prompts coordination and balanced participation). As access to research participants is not yet available, I release an open synthetic benchmark dataset that mirrors ASTRA’s logging schema and a prespecified between-subjects design ( participants; 360 sessions; 1440 task episodes) across a bank of 20 short Python programming tasks. The dataset includes turn-level dialogue traces and task-level artefacts designed to support log-operational research questions about interaction dynamics, participation balance and reciprocal engagement in dyads, and performance and verification behaviours. Descriptive summaries and illustrative models indicate that the benchmark yields measurable condition-differentiated patterns consistent with the simulation assumptions. I emphasise that these findings are simulated evidence intended for benchmarking, measurement feasibility, and reproducible pipeline development, not causal estimates of learning effects, while providing a transparent analysis blueprint for future ethics-approved validation studies.