Much Gracias: Semi-supervised Code-switch Detection for Spanish-English: How far can we get?
- Dana-Maria Iliescu,
- Rasmus Grand,
- ,
- Sara Qirko
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 65 (6 pages)Publication milestones
- Published - 06/2021
Publication status
Published - 06/2021
Publisher
Association for Computational Linguistics, United StatesBook series
- Book series name: Proceedings of the Fifth Workshop on Computational Approaches to Linguistic Code-Switching
Host publication title
Proceedings of the Fifth Workshop on Computational Approaches to Linguistic Code-SwitchingAbstract
Because of globalization, it is becoming more and more common to use multiple languages in a single utterance, also called code-switching. This results in special linguistic structures and, therefore, poses many challenges for Natural Language Processing. Existing models for language identification in code-switched data are all supervised, requiring annotated training data which is only available for a limited number of language pairs. In this paper, we explore semi-supervised approaches, that exploit out-of-domain mono-lingual training data. We experiment with word uni-grams, word n-grams, character n-grams, Viterbi Decoding, Latent Dirichlet Allocation, Support Vector Machine and Logistic Regression. The Viterbi model was the best semi-supervised model, scoring a weighted F1 score of 92.23%, whereas a fully supervised state-of-the-art BERT-based model scored 98.43%.
Access to documents
Final published version, 299.81 KB
License:CC BY-NC-SA, opens in new tab
Final published version
License:CC BY-NC-SA, opens in new tab
Related Event
Title
The 5th Workshop on Computational Approaches to Linguistic Code-Switching
