Skip to search boxSkip to navigationSkip to main content

Dissecting Biases in Relation Extraction: A Cross-Dataset Analysis on People’s Gender and Origin

  • Marco Antonio Stranisci
    ,
  • Pere-Lluís Huguet Cabot
    ,
  • ,
  • Roberto Navigli
  • University of Turin
    ,
  • Sapienza University of Rome
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 190-202

Publication milestones

  • Published - 2024

Publication status

Published - 2024

Publication IDs

  • Scopus: 85204359185

Host publication title

Proceedings of the 5th Workshop on Gender Bias in Natural Language Processing (GeBNLP)

Abstract

Relation Extraction (RE) is at the core of many Natural Language Understanding tasks, including knowledge-base population and Question Answering. However, any Natural Language Processing system is exposed to biases, and the analysis of these has not received much attention in RE. We propose a new method for inspecting bias in the RE pipeline, which is completely transparent in terms of interpretability. Specifically, in this work we analyze biases related to gender and place of birth. Our methodology includes (i) obtaining semantic triplets (subject, object, semantic relation) involving ‘person’ entities from RE resources, (ii) collecting meta-information (‘gender’ and ‘place of birth’) using Entity Linking technologies, and then (iii) analyze the distribution of triplets across different groups (e.g., men versus women). We investigate bias at two levels: In the training data of three commonly used RE datasets (SREDFM, CrossRE, NYT), and in the predictions of a state-of-the-art RE approach (ReLiK). To enable cross-dataset analysis, we introduce a taxonomy of relation types mapping the label sets of different RE datasets to a unified label space. Our findings reveal that bias is a compounded issue affecting underrepresented groups within data and predictions for RE.

Publication metrics

PlumX

Citations
3

Access to documents

Related Event

Title

Workshop on Gender Bias in Natural Language Processing

Event type

Conference

Degree of recognition

International event

Date

16/08/2024

Location

BangkokThailand