NLP North at WNUT-2020 Task 2: Pre-training versus Ensembling for Detection of Informative COVID-19 English Tweets
- Anders Giovanni Møller,
- ,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 331-336Publication milestones
- Published - 11/2020
Publication status
Published - 11/2020
Publisher
Association for Computational Linguistics, United StatesHost publication title
Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020)Abstract
With the COVID-19 pandemic raging world-wide since the beginning of the 2020 decade,the need for monitoring systems to track relevant information on social media is vitally important. This paper describes our submission to the WNUT-2020 Task 2: Identification of informative COVID-19 English Tweets. We investigate the effectiveness for a variety of classification models, and found that domain-specific pre-trained BERT models lead to the best performance. On top of this, we attempt a variety of ensembling strategies, but these at-tempts did not lead to further improvements.Our final best model, the standalone CT-BERT model, proved to be highly competitive, leading to a shared first place in the shared task.Our results emphasize the importance of do-main and task-related pre-training.
Access to documents
Accepted author manuscript, 162.84 KB
Accepted author manuscript
Related Event
Title
The Sixth Workshop on Noisy User-generated Text
Event type
ConferenceDate
19/11/2020 - 19/11/2020Location
Online
