Finding the needle in a haystack: Extraction of Informative COVID-19 Danish Tweets
- Benjamin Ahrentløv Olsen,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 11–19Publication milestones
- Published - 2021
Publication status
Published - 2021
Publisher
Association for Computational Linguistics, United StatesHost publication title
Proceedings of the 2021 EMNLP Workshop W-NUT: The Seventh Workshop on Noisy User-generated TextAbstract
Finding informative COVID-19 posts in a stream of tweets is very useful to monitor health-related updates. Prior work focused on a balanced data setup and on English, but in- formative tweets are rare, and English is only one of the many languages spoken in the world. In this work, we introduce a new dataset of 5,000 tweets for finding informative COVID- 19 tweets for Danish. In contrast to prior work, which balances the label distribution, we model the problem by keeping its natural dis- tribution. We examine how well a simple prob- abilistic model and a convolutional neural net- work (CNN) perform on this task. We find a weighted CNN to work well but it is sensi- tive to embedding and hyperparameter choices. We hope the contributed dataset is a starting point for further work in this direction.
Access to documents
Accepted author manuscript, 987.91 KB
Accepted author manuscript
Related Event
Title
The Seventh Workshop on Noisy User-generated Text
Event type
ConferenceDate
11/11/2021 - 11/11/2021Location
VIRTUAL
