Skip to search boxSkip to navigationSkip to main content

Are you sure? Measuring models bias in content moderation through uncertainty

  • Northeastern University
    ,
  • Universitá del Piemonte Orientale
    ,
  • Heriot-Watt University
    ,
  • University of Torino
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 18061–18076 (16 pages)

Publication milestones

  • Published - 2025

Publication status

Published - 2025

Publisher

Association for Computational Linguistics, United States
9798891763357, 9798891763357, 9798891763357

Publication IDs

  • ORCID: /0000-0001-9337-7250/work/220441844
  • Scopus: 105028970175

Host publication title

Emnlp 2025 2025 Conference on Empirical Methods in Natural Language Processing Findings of Emnlp 2025

Abstract

Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several resources and benchmark corpora have been developed to challenge this issue, measuring the fairness of models in content moderation remains an open issue. In this work, we present an unsupervised approach that benchmarks models on the basis of their uncertainty in classifying messages annotated by people belonging to vulnerable groups. We use uncertainty, computed by means of the conformal prediction technique, as a proxy to analyze the bias of 11 models (LMs and LLMs) against women and non-white annotators and observe to what extent it diverges from metrics based on performance, such as the F1 score. The results show that some pre-trained models predict with high accuracy the labels coming from minority groups, even if the confidence in their prediction is low. Therefore, by measuring the confidence of models, we are able to see which groups of annotators are better represented in pre-trained models and lead the debiasing process of these models before their effective use.

Publication metrics

Funding Details

The work of S. Frenda is supported by the EPSRC project “Equally Safe Online” (EP/W025493/1).
FundersFunding numbers
UKRI - UK Research and Innovation
EP/W025493/1

Related Event

Title

2025 Conference on Empirical Methods in Natural Language Processing

Event type

Conference

Degree of recognition

International event

Date

04/11/2024 - 09/11/2024

Location

SuzhouChina