TY - GEN
T1 - From Data to Model in Bias
T2 - 19th ACM International Conference on Web Search and Data Mining, WSDM 2026
AU - You, Jaebeom
AU - Lee, Jaewon
AU - Lee, Sehun
AU - Kwon, Hyuk Yoon
N1 - Publisher Copyright:
© 2026 Owner/Author.
PY - 2026/2/21
Y1 - 2026/2/21
N2 - While political biases in Large Language Models (LLMs) have been consistently reported, the underlying question of whether such biases stem from systematic distributions in pretraining data remains insufficiently examined. Existing work has largely focused on post-hoc bias mitigation, with limited content-level analysis in large-scale web corpora. This study presents the first comprehensive statistical analysis of political bias in a large-scale web corpus, C4, a widely used web corpus that underpins major LLMs such as T5, PaLM, and LaMDA. We analyze 15 politically sensitive topics using a multi-perspective, persona-based LLM annotation framework, combined with rigorous statistical validation through equivalence testing and multiple testing correction. Our findings reveal clear patterns of left-leaning and supportive bias in C4 across two dimensions: political orientation and topic-specific stance. In particular, strong progressive leanings are observed in social and cultural domains (e.g., gender equality, LGBTQ rights, abortion rights), while economic topics exhibit relatively balanced distributions. To explore data-to-model bias transfer, we conduct experiments examining correlations between corpus-level bias and LLM responses, and evaluate the effect of fine-tuning on model behavior. The results provide evidence that political bias in web corpora can propagate into model outputs. These findings highlight the importance of dataset-level bias analysis in understanding and mitigating political bias in LLMs. We argue that responsible AI development must incorporate systematic data curation practices. All source code and scripts used in this study are publicly available at: https://anonymous.4open.science/r/C4_analysis-D7C2.
AB - While political biases in Large Language Models (LLMs) have been consistently reported, the underlying question of whether such biases stem from systematic distributions in pretraining data remains insufficiently examined. Existing work has largely focused on post-hoc bias mitigation, with limited content-level analysis in large-scale web corpora. This study presents the first comprehensive statistical analysis of political bias in a large-scale web corpus, C4, a widely used web corpus that underpins major LLMs such as T5, PaLM, and LaMDA. We analyze 15 politically sensitive topics using a multi-perspective, persona-based LLM annotation framework, combined with rigorous statistical validation through equivalence testing and multiple testing correction. Our findings reveal clear patterns of left-leaning and supportive bias in C4 across two dimensions: political orientation and topic-specific stance. In particular, strong progressive leanings are observed in social and cultural domains (e.g., gender equality, LGBTQ rights, abortion rights), while economic topics exhibit relatively balanced distributions. To explore data-to-model bias transfer, we conduct experiments examining correlations between corpus-level bias and LLM responses, and evaluate the effect of fine-tuning on model behavior. The results provide evidence that political bias in web corpora can propagate into model outputs. These findings highlight the importance of dataset-level bias analysis in understanding and mitigating political bias in LLMs. We argue that responsible AI development must incorporate systematic data curation practices. All source code and scripts used in this study are publicly available at: https://anonymous.4open.science/r/C4_analysis-D7C2.
KW - context-aware data scrap-ing
KW - llm-generated persona
KW - search engines
KW - statistical significance testing
UR - https://www.scopus.com/pages/publications/105033150666
U2 - 10.1145/3773966.3777990
DO - 10.1145/3773966.3777990
M3 - Conference contribution
AN - SCOPUS:105033150666
T3 - WSDM 2026 - Proceedings of the 19th ACM International Conference on Web Search and Data Mining
SP - 860
EP - 870
BT - WSDM 2026 - Proceedings of the 19th ACM International Conference on Web Search and Data Mining
PB - Association for Computing Machinery, Inc
Y2 - 22 February 2026 through 26 February 2026
ER -