Skip to search boxSkip to navigationSkip to main content

ESCOXLM-R: Multilingual Taxonomy-driven Pre-training for the Job Market Domain

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 11871–11890 (20 pages)

Publication milestones

  • Published - 07/2023

Publication status

Published - 07/2023

Place of publication

Toronto, Canada

Volume

Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Publisher

Association for Computational Linguistics, United States

Publication IDs

  • Scopus: 85174422231

Host publication title

The 61st Annual Meeting of the Association for Computational Linguistics

Abstract

The increasing number of benchmarks for Natural Language Processing (NLP) tasks in the computational job market domain highlights the demand for methods that can handle job-related tasks such as skill extraction, skill classification, job title classification, and de-identification. While some approaches have been developed that are specific to the job market domain, there is a lack of generalized, multilingual models and benchmarks for these tasks. In this study, we introduce a language model called ESCOXLM-R, based on XLM-R-large, which uses domain-adaptive pre-training on the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy, covering 27 languages. The pre-training objectives for ESCOXLM-R include dynamic masked language modeling and a novel additional objective for inducing multilingual taxonomical ESCO relations. We comprehensively evaluate the performance of ESCOXLM-R on 6 sequence labeling and 3 classification tasks in 4 languages and find that it achieves state-of-the-art results on 6 out of 9 datasets. Our analysis reveals that ESCOXLM-R performs better on short spans and outperforms XLM-R-large on entity-level and surface-level span-F1, likely due to ESCO containing short skill and occupation titles, and encoding information on the entity-level.

Publication metrics

PlumX, opens in new tab

Citations
28
Captures
40

Related Event

Title

Annual Meeting of the Association for Computational Linguistics

Event type

Conference

Degree of recognition

International event

Date

09/07/2023 - 14/07/2023

Location

TorontoCanada