CL-MoNoise: Cross-lingual Lexical Normalization
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 510 (4 pages)Publication milestones
- Published - 10/2021
Publication status
Published - 10/2021
Publisher
Association for Computational Linguistics, United StatesHost publication title
Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021)Abstract
Social media is notoriously difficult to process for existing natural language processing tools, because of spelling errors, non-standard words, shortenings, non-standard capitalization and punctuation. One method to circumvent these issues is to normalize input data before processing. Most previous work has focused on only one language, which is mostly English. In this paper, we are the first to propose a model for cross-lingual normalization, with which we participate in the WNUT 2021 shared task. To this end, we use MoNoise as a starting point, and make a simple adaptation for cross-lingual application. Our proposed model outperforms the leave-as-is baseline provided by the organizers which copies the input. Furthermore, we explore a completely different model which converts the task to a sequence labeling task. Performance of this second system is low, as it does not take capitalization into account in our implementation.
Access to documents
Final published version
License:CC BY-NC-SA, opens in new tab
Related Event
Title
Workshop on Noisy User-generated Text
Event type
WorkshopDegree of recognition
International eventDate
11/11/2021 - 11/11/2021Location
VIRTUAL
