Are you sure? Measuring models bias in content moderation through uncertainty
- Alessandra Urbinati,
- Mirko Lai,
- Simona Frenda,
- Northeastern University,
- Universitá del Piemonte Orientale,
- Heriot-Watt University,
- University of Torino
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 18061–18076 (16 pages)Publication milestones
- Published - 2025
Publication status
Published - 2025
Publisher
Association for Computational Linguistics, United StatesISBN (Print)
9798891763357, 9798891763357, 9798891763357Publication IDs
- ORCID: /0000-0001-9337-7250/work/220441844
- Scopus: 105028970175
Host publication title
Emnlp 2025 2025 Conference on Empirical Methods in Natural Language Processing Findings of Emnlp 2025Abstract
Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several resources and benchmark corpora have been developed to challenge this issue, measuring the fairness of models in content moderation remains an open issue. In this work, we present an unsupervised approach that benchmarks models on the basis of their uncertainty in classifying messages annotated by people belonging to vulnerable groups. We use uncertainty, computed by means of the conformal prediction technique, as a proxy to analyze the bias of 11 models (LMs and LLMs) against women and non-white annotators and observe to what extent it diverges from metrics based on performance, such as the F1 score. The results show that some pre-trained models predict with high accuracy the labels coming from minority groups, even if the confidence in their prediction is low. Therefore, by measuring the confidence of models, we are able to see which groups of annotators are better represented in pre-trained models and lead the debiasing process of these models before their effective use.
Publication metrics
PlumX, opens in new tab
Captures
3
Funding Details
The work of S. Frenda is supported by the EPSRC project “Equally Safe Online” (EP/W025493/1).
FundersFunding numbers
UKRI - UK Research and Innovation
EP/W025493/1
Access to documents
Related Event
Title
2025 Conference on Empirical Methods in Natural Language Processing
Event type
ConferenceDegree of recognition
International eventDate
04/11/2024 - 09/11/2024Location
SuzhouChina
