Skip to search boxSkip to navigationSkip to main content

DanTok: Domain Beats Language for Danish Social Media POS Tagging

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 271–279 (9 pages)

Publication milestones

  • Published - 2023

Publication status

Published - 2023

Publication IDs

  • Scopus: 85217079230

Host publication title

Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)

Abstract

Language from social media remains challenging to process automatically, especially for non-English languages. In this work, we introduce the first NLP dataset for TikTok comments and the first Danish social media dataset with part-of-speech annotation. We further supply annotations for normalization, code-switching, and annotator uncertainty. As transferring models to such a highly specialized domain is non-trivial, we conduct an extensive study into which source data and modeling decisions most impact the performance. Surprisingly, transferring from in-domain data, even from a different language, outperforms in-language, out-of-domain training. These benefits nonetheless rely on the underlying language models having been at least partially pre-trained on data from the target language. Using our additional annotation layers, we further analyze how normalization, code-switching, and human uncertainty affect the tagging accuracy.

Publication metrics

PlumX

Captures
19
Citations
1

Related Event

Title

Nordic Conference on Computational Linguistics

Event type

Conference

Degree of recognition

International event

Date

22/05/2023 - 24/05/2023

Location

TórshavnFaroe Islands