Skip to search boxSkip to navigationSkip to main content

A Dataset for the Detection of Dehumanizing Language

Research Output:
Other contribution
Other contribution

Open access

Publication Information

Output type

Research Output:
Other contribution
Other contribution

Original language

English

Publication milestones

  • Published - 2024

Publication status

Published - 2024

Publication IDs

  • ORCID: /0000-0002-6103-7275/work/156308212
  • Scopus: 85185868622

Abstract

Dehumanization is a mental process that enables the exclusion and ill treatment of a group of people. In this paper, we present two data sets of dehumanizing text, a large, automatically collected corpus and a smaller, manually annotated data set. Both data sets include a combination of political discourse and dialogue from movie subtitles. Our methods give us a broad and varied amount of dehumanization data to work with, enabling further exploratory analysis and automatic classification of dehumanization patterns