Discovering Aspectual Classes of Russian Verbs in Untagged Large Corpora
- Aleksandr Drozd,
- ,
- Satoshi Matsuoka
- Tokyo Institute of Technology
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewPublication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 61-68 (8 pages)Publication milestones
- Published - 2015
Publication status
Published - 2015
Publication IDs
- Scopus: 84964507906
Host publication title
Proceedings of 2015 IEEE International Conference on Data Science and Data Intensive Systems (DSDIS)Abstract
This paper presents a case study of discovering and classifying verbs in large web-corpora. Many tasks in natural language processing require corpora containing billions of words, and with such volumes of data co-occurrence extraction becomes one of the performance bottlenecks in the Vector Space Models of computational linguistics. We propose a co-occurrence extraction kernel based on ternary trees as an alternative (or a complimentary stage) to conventional map-reduce based approach, this kernel achieves an order of magnitude improvement in memory footprint and processing speed. Our classifier successfully and efficiently identified verbs in a 1.2-billion words untagged corpus of Russian fiction and distinguished between their two aspectual classes. The model proved efficient even for low-frequency vocabulary, including nonce verbs and neologisms.
Publication metrics
PlumX, opens in new tab
Captures
9
Citations
4
