SHR++: An Interface for Morpho-syntactic annotation of Sanskrit Corpora
- Amrith Krishna,
- Shiv Vidhyut,
- Dilpreet Chawla,
- Sruti Sambhavi,
- Pawan Goyal
- ,
- ,
- International Institute of Information Technology Bangalore,
- NIT - National Institute of Technology Rourkela,
- Indian Institute of Technology Kharagpur
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 7069–7076Publication milestones
- Published - 02/2020
Publication status
Published - 02/2020
Publisher
Association for Computational Linguistics, United StatesPublication IDs
- Scopus: 85096512176
Host publication title
Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020),Abstract
We propose a web-based annotation framework, SHR++, for morpho-syntactic annotation of corpora in Sanskrit. SHR++ is designed
to generate annotations for the word-segmentation, morphological parsing and dependency analysis tasks in Sanskrit. It incorporates
analyses and predictions from various tools designed for processing texts in Sanskrit, and utilises them to ease the cognitive load of
the human annotators. Specifically, SHR++ uses Sanskrit Heritage Reader (Goyal and Huet, 2016), a lexicon driven shallow parser
for enumerating all the phonetically and lexically valid word splits along with their morphological analyses for a given string. This
would help the annotators in choosing the solutions, rather than performing the segmentations by themselves. Further, predictions from
a word segmentation tool (Krishna et al., 2018) are added as suggestions that can aid the human annotators in their decision making.
Our evaluation shows that enabling this segmentation suggestion component reduces the annotation time by 20.15 %. SHR++ can
be accessed online at http://vidhyut97.pythonanywhere.com/ and the codebase, for the independent deployment of the system elsewhere, is hosted at https://github.com/iamdsc/smart-sanskrit-annotator.
to generate annotations for the word-segmentation, morphological parsing and dependency analysis tasks in Sanskrit. It incorporates
analyses and predictions from various tools designed for processing texts in Sanskrit, and utilises them to ease the cognitive load of
the human annotators. Specifically, SHR++ uses Sanskrit Heritage Reader (Goyal and Huet, 2016), a lexicon driven shallow parser
for enumerating all the phonetically and lexically valid word splits along with their morphological analyses for a given string. This
would help the annotators in choosing the solutions, rather than performing the segmentations by themselves. Further, predictions from
a word segmentation tool (Krishna et al., 2018) are added as suggestions that can aid the human annotators in their decision making.
Our evaluation shows that enabling this segmentation suggestion component reduces the annotation time by 20.15 %. SHR++ can
be accessed online at http://vidhyut97.pythonanywhere.com/ and the codebase, for the independent deployment of the system elsewhere, is hosted at https://github.com/iamdsc/smart-sanskrit-annotator.
Publication metrics
PlumX
Citations
5
Captures
66
Access to documents
Final published version, 1.21 MB
License:CC BY-NC, opens in new tab
Final published version
License:CC BY-NC, opens in new tab
Related Event
Title
Conference on Linguistic Resources and Evaluation
Event type
ConferenceDegree of recognition
International eventDate
20/06/2022 - 25/06/2022Location
Palais du PharoMarseilleFrance
