Skip to search boxSkip to navigationSkip to main content

Docent: A Document-Level Decoder for Phrase-Based Statistical Machine Translation

  • Uppsala University
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Publication milestones

  • Published - 09/08/2013

Publication status

Published - 09/08/2013

Publication IDs

  • ORCID: /0000-0002-6103-7275/work/106363244

Host publication title

Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics: System Demonstrations

Abstract

We describe Docent, an open-source decoder for statistical machine translation that breaks with the usual sentence-by-sentence paradigm and translates complete documents as units. By taking translation to the document level, our decoder can handle feature models with arbitrary discourse-wide dependencies and constitutes an essential infrastructure component in the quest for discourse-aware SMT models.