DanTok: Domain Beats Language for Danish Social Media POS Tagging
- Kia Kirstein Hansen,
- Maria Jung Barrett,
- ,
- Cathrine Damgaard,
- Trine Eriksen,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 271–279 (9 pages)Publication milestones
- Published - 2023
Publication status
Published - 2023
Publication IDs
- Scopus: 85217079230
Host publication title
Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)Abstract
Language from social media remains challenging to process automatically, especially for non-English languages. In this work, we introduce the first NLP dataset for TikTok comments and the first Danish social media dataset with part-of-speech annotation. We further supply annotations for normalization, code-switching, and annotator uncertainty. As transferring models to such a highly specialized domain is non-trivial, we conduct an extensive study into which source data and modeling decisions most impact the performance. Surprisingly, transferring from in-domain data, even from a different language, outperforms in-language, out-of-domain training. These benefits nonetheless rely on the underlying language models having been at least partially pre-trained on data from the target language. Using our additional annotation layers, we further analyze how normalization, code-switching, and human uncertainty affect the tagging accuracy.
Publication metrics
PlumX
Captures
19
Citations
1
Access to documents
Final published version
Related Event
Title
Nordic Conference on Computational Linguistics
Event type
ConferenceDegree of recognition
International eventDate
22/05/2023 - 24/05/2023Location
TórshavnFaroe Islands
