Skip to search boxSkip to navigationSkip to main content

A Neural Model for Part-of-Speech Tagging in Historical Texts

  • Uppsala University
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Publication milestones

  • Published - 16/12/2016

Publication status

Published - 16/12/2016
978-4-87974-702-0

Publication IDs

  • ORCID: /0000-0002-6103-7275/work/106363235
  • Scopus: 85055000572

Host publication title

Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers

Abstract

Historical texts are challenging for natural language processing because they differ linguistically from modern texts and because of their lack of orthographical and grammatical standardisation. We use a character-level neural network to build a part-of-speech (POS) tagger that can process historical data directly without requiring a separate spelling normalisation stage. Its performance in a Swedish verb identification and a German POS tagging task is similar to that of a two-stage model. We analyse the performance of this tagger and a more traditional baseline system, discuss some of the remaining problems for tagging historical data and suggest how the flexibility of our neural tagger could be exploited to address diachronic divergences in morphology and syntax in early modern Swedish with the help of data from closely related languages.

Publication metrics

PlumX

Citations
8
Captures
74

Access to documents

Final published version
License:Unspecified