Skip to search boxSkip to navigationSkip to main content

When Simple n-gram Models Outperform Syntactic Approaches: Discriminating between Dutch and Flemish

  • Leiden University
    ,
  • University of Groningen
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 244-253

Publication milestones

  • Published - 2018

Publication status

Published - 2018

Publisher

Association for Computational Linguistics, United States
978-1-948087-55-1

Publication IDs

  • Scopus: 85122271504

Host publication title

Proceedings of the Fifth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial)

Abstract

In this paper we present the results of our participation in the Discriminating between Dutch and Flemish in Subtitles VarDial 2018 shared task. We try techniques proven to work well for discriminating between language varieties as well as explore the potential of using syntactic features, i.e. hierarchical syntactic subtrees. We experiment with different combinations of features. Discriminating between these two languages turned out to be a very hard task, not only for a machine: human performance is only around 0.51 F1 score; our best system is still a simple Naive Bayes model with word unigrams and bigrams. The system achieved an F1 score (macro)
of 0.62, which ranked us 4th in the shared task.

Publication metrics

PlumX

Citations
6
Captures
70