Skip to search boxSkip to navigationSkip to main content

Predictive Text Entry for Agglutinative Languages Using Unsupervised Morphological Segmentation

  • Miikka Silfverberg
    ,
  • Krister Lindén
    ,
  • Hissu Hyvärinen
  • University of Helsinki
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Host publication Subtitle

CICLing 2012: Computational Linguistics and Intelligent Text Processing

Original language

English

Pages from-to (Number of pages)

Pages 478-489

Publication milestones

  • Published - 2012

Publication status

Published - 2012

Publisher

Springer, United States, Germany

Book series

  • Book series name: Lecture Notes in Computer Science
    Volume: 7182
    ISSN: 0302-9743
978-3-642-28600-1

ISBN (Electronic)

978-3-642-28601-8

Publication IDs

  • Scopus: 84858314006

Host publication title

International Conference on Intelligent Text Processing and Computational Linguistics

Abstract

Systems for predictive text entry on ambiguous keyboards typically rely on dictionaries with word frequencies which are used to suggest the most likely words matching user input. This approach is insufficient for agglutinative languages, where morphological phenomena increase the rate of out-of-vocabulary words. We propose a method for text entry, which circumvents the problem of out-of-vocabulary words, by replacing the dictionary with a Markov chain on morph sequences combined with a third order hidden Markov model (HMM) mapping key sequences to letter sequences and phonological constraints for pruning suggestion lists. We evaluate our method by constructing text entry systems for Finnish and Turkish and comparing our systems with published text entry systems and the text entry systems of three commercially available mobile phones. Measured using the keystrokes per character ratio (KPC) [8], we achieve superior results. For training, we use corpora, which are segmented using unsupervised morphological segmentation.

Publication metrics

PlumX, opens in new tab

Captures
2
Citations
3