Skip to search boxSkip to navigationSkip to main content

De-identification of Privacy-related Entities in Job Postings

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 210-221

Publication milestones

  • Accepted/In press - 22/03/2021
  • Published - 21/05/2021

Publication status

Published - 21/05/2021

Publisher

Association for Computational Linguistics, United States

Book series

  • Book series name: Linköping Electronic Conference Proceedings
    Series number: 21
    Volume: 178

Host publication title

Proceedings of the 23rd Nordic Conference on Computational Linguistics

Abstract

De-identification is the task of detecting privacy-related entities in text, such as person names, emails and contact data. It has been well-studied within the medical domain. The need for de-identification technology is increasing, as privacy-preserving data handling is in high demand in many domains. In this paper, we focus on job postings. We present JobStack, a new corpus for de-identification of personal data in job vacancies on Stackoverflow. We introduce baselines, comparing Long-Short Term Memory (LSTM) and Transformer models. To improve upon these baselines, we experiment with contextualized embeddings and distantly related auxiliary data via multi-task learning. Our results show that auxiliary data improves de-identification performance.

Access to documents

Accepted author manuscript, 204.71 KB

Related Event

Title

Nordic Conference on Computational Linguistics

Event type

Conference

Degree of recognition

International event

Date

31/05/2021 - 02/06/2021

Location

RejkjavikIceland