Discriminating Between Similar Nordic Languages
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 67–75Publication milestones
- Published - 20/04/2021
Publication status
Published - 20/04/2021
Publisher
Association for Computational Linguistics, United StatesHost publication title
Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and DialectsAbstract
Automatic language identification is a challenging problem. Discriminating between closely related languages is especially difficult. This paper presents a machine learning approach for automatic language identification for the Nordic languages, which often suffer miscategorisation by existing state-of-the-art tools. Concretely we will focus on discrimination between six Nordic languages: Danish, Swedish, Norwegian (Nynorsk), Norwegian (Bokmål), Faroese and Icelandic.
Access to documents
Final published version
License:CC BY, opens in new tab
Related Event
Title
Workshop on NLP for Similar Languages, Varieties and Dialects
Event type
WorkshopDate
20/04/2021 - 20/04/2021Location
VIRTUAL
