DAN+: Danish Nested Named Entities and Lexical Normalization
- ,
- Kristian Nørgaard Jensen,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 6649–6662Publication milestones
- Published - 12/2020
Publication status
Published - 12/2020
Publisher
Association for Computational Linguistics, United StatesHost publication title
The 28th International Conference on Computational LinguisticsAbstract
This paper introduces DAN+, a new multi-domain corpus and annotation guidelines for Dan- ish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language. We empirically assess three strategies to model the two-layer Named Entity Recognition (NER) task. We compare transfer capabilities from German versus in-language annotation from scratch. We examine language-specific versus multilingual BERT, and study the effect of lexical normalization on NER. Our results show that 1) the most robust strategy is multi-task learning which is rivaled by multi-label decoding, 2) BERT-based NER models are sensitive to domain shifts, and 3) in-language BERT and lexical normalization are the most beneficial on the least canonical data. Our results also show that an out-of-domain setup remains challenging, while performance on news plateaus quickly. This highlights the importance of cross-domain evaluation of cross-lingual transfer.
Access to documents
Accepted author manuscript, 170.01 KB
Related Event
Title
International Conference on Computational Linguistics
Event type
ConferenceDegree of recognition
International eventDate
08/12/2020 - 13/12/2020Location
BarcelonaSpain
