llm-bias
Filtering by topic llm-bias(2)Clear all filters
- PaperarXiv — AI in Education (cs.CY)13 Jul 2026
The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
Alexis Popovici, Andrei Ionascu, Adrian-Marius Dumitran
A systematic audit of four LLMs acting as history tutors found that safety-aligned models exhibit epistemic paternalism, differentially refusing 76.7% of educational requests from low-tier students and reducing access to complex geopolitical content for marginalized learners. The study identifies patterns including differential refusal, epistemic gatekeeping, agency theft, and elite hermeneutics, arguing that current safety alignment functions as a paternalistic filter that perpetuates narrative segregation.
Original abstract
As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regarding the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. We uncover four interconnected patterns of \emph{epistemic paternalism}: (1)~\textbf{Differential Refusal}, where safety-aligned models block 76.7\% of educational requests from low-tier students; (2)~\textbf{Epistemic Gatekeeping}, evidenced by a 3$\times$ reduction in access to geopolitical complexity (e.g., the contested ``coup theory'') for marginalized learners; (3)~\textbf{Agency Theft}, a lexical shift where models like LLaMA produce a 5$\times$ higher victimization-to-politics vocabulary ratio for Roma students compared to elite peers; and (4)~\textbf{Elite Hermeneutics}, where AI tutors disproportionately withhold epistemic confidence and justification scores from low-resource demographic profiles. We argue that current safety alignment acts as a paternalistic filter, transforming conversational AI into agents of narrative segregation -- a manifestation of \emph{hermeneutical injustice} in Fricker's~\cite{fricker2007} sense that demands urgent pedagogical auditing.
- PaperETS Research Report Series20 Apr 2026
On the Representation of Racial and Ethnic Subgroups in AI-generated Texts: A Case Study in Automated Essay Scoring
Akshay Badola, Mo Zhang, Chen Li
Using GPT-4 and GPT-4o to generate essays for specific racial/ethnic subgroups from example essays, this study finds that the generated racial distribution does not match the real distribution. Augmenting automated essay scoring training data with these LLM-generated essays, even when race is mispredicted, reduces bias without harming performance.
Original abstract
In this study, we assess the capability of LLMs in generating essays of a specific race/ethnicity after being given example essays and rubric, and investigate the efficacy of data augmented in this manner for Automated Essay Scoring with respect to model performance and bias. In a series of experiments, we use models GPT-4 and GPT-4o, and ask them to generate essays from a given subgroup after inferring the race/ethnicity of the writer. We find that while LLMs can be directed to generate essays for specific demographic groups, the inferred racial and ethnic distribution in the generated data does not closely mirror the actual distribution observed in the source dataset. We augment existing data for underrepresented subgroups with LLM generated data separated into two groups with correct LLM race prediction and with incorrect race prediction and assess the improvement in agreement with human scores with quadratic weighted Kappa and bias mitigation as change in standardized mean difference. Our analysis shows that while LLMs struggle to predict the race accurately from given samples, augmentation with such data can be helpful to mitigate bias regardless. Suggested citation: Badola, A., Zhang, M., & Li, Chen. (in press). On the representation of racial and ethnic subgroups in AI-generated texts: A case study in automated essay scoring. ETS Research Report Series. https://doi.org/10.64634/ac01td58