Skip to main navigation Skip to search Skip to main content

From Data to Model in Bias: A Statistical Analysis of Political Bias in the C4 Corpus and Its Impact on LLMs

  • Seoul National University of Science and Technology (SNUST)
  • Kookmin University
  • Korea University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

While political biases in Large Language Models (LLMs) have been consistently reported, the underlying question of whether such biases stem from systematic distributions in pretraining data remains insufficiently examined. Existing work has largely focused on post-hoc bias mitigation, with limited content-level analysis in large-scale web corpora. This study presents the first comprehensive statistical analysis of political bias in a large-scale web corpus, C4, a widely used web corpus that underpins major LLMs such as T5, PaLM, and LaMDA. We analyze 15 politically sensitive topics using a multi-perspective, persona-based LLM annotation framework, combined with rigorous statistical validation through equivalence testing and multiple testing correction. Our findings reveal clear patterns of left-leaning and supportive bias in C4 across two dimensions: political orientation and topic-specific stance. In particular, strong progressive leanings are observed in social and cultural domains (e.g., gender equality, LGBTQ rights, abortion rights), while economic topics exhibit relatively balanced distributions. To explore data-to-model bias transfer, we conduct experiments examining correlations between corpus-level bias and LLM responses, and evaluate the effect of fine-tuning on model behavior. The results provide evidence that political bias in web corpora can propagate into model outputs. These findings highlight the importance of dataset-level bias analysis in understanding and mitigating political bias in LLMs. We argue that responsible AI development must incorporate systematic data curation practices. All source code and scripts used in this study are publicly available at: https://anonymous.4open.science/r/C4_analysis-D7C2.

Original languageEnglish
Title of host publicationWSDM 2026 - Proceedings of the 19th ACM International Conference on Web Search and Data Mining
PublisherAssociation for Computing Machinery, Inc
Pages860-870
Number of pages11
ISBN (Electronic)9798400722929
DOIs
StatePublished - 21 Feb 2026
Event19th ACM International Conference on Web Search and Data Mining, WSDM 2026 - Boise, United States
Duration: 22 Feb 202626 Feb 2026

Publication series

NameWSDM 2026 - Proceedings of the 19th ACM International Conference on Web Search and Data Mining

Conference

Conference19th ACM International Conference on Web Search and Data Mining, WSDM 2026
Country/TerritoryUnited States
CityBoise
Period22/02/2626/02/26

Keywords

  • context-aware data scrap-ing
  • llm-generated persona
  • search engines
  • statistical significance testing

Fingerprint

Dive into the research topics of 'From Data to Model in Bias: A Statistical Analysis of Political Bias in the C4 Corpus and Its Impact on LLMs'. Together they form a unique fingerprint.

Cite this