Skip to search boxSkip to navigationSkip to main content

Challenges in Annotating and Parsing Spoken, Code-switched, Frisian-Dutch Data

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 50-58

Publication milestones

  • Published - 04/2021

Publication status

Published - 04/2021

Publisher

Association for Computational Linguistics, United States

Publication IDs

  • Scopus: 85104376384

Host publication title

Proceedings of the Second Workshop on Domain Adaptation for NLP

Abstract

While high performance have been obtained for high-resource languages, performance on low-resource languages lags behind. In this paper we focus on the parsing of the low-resource language Frisian. We use a sample of code-switched, spontaneously spoken data, which proves to be a challenging setup. We propose to train a parser specifically tailored towards the target domain, by selecting instances from multiple treebanks. Specifically, we use Latent Dirichlet Allocation (LDA), with word and character N-grams. We use a deep biaffine parser initialized with mBERT. The best single source treebank (nl_alpino) resulted in an LAS of 54.7 whereas our data selection outperformed the single best transfer treebank and led to 55.6 LAS on the test data. Additional experiments consisted of removing diacritics from our Frisian data, creating more similar training data by cropping sentences and running our best model using XLM-R. These experiments did not lead to a better performance.

Publication metrics

PlumX

Captures
58
Citations
17

Related Event

Title

Workshop on Domain Adaptation for NLP

Description

Workshop held at EACL conference

Event type

Workshop

Degree of recognition

International event

Date

20/04/2021

Location

KyivUkraine