Identifying Open Challenges in Language Identification
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 18207-18227 (21 pages)Publication milestones
- Published - 01/07/2025
Publication status
Published - 01/07/2025
Place of publication
Vienna, AustriaPublisher
Association for Computational Linguistics, United StatesISBN (Print)
979-8-89176-251-0Publication IDs
- Scopus: 105021034285
Host publication title
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)Host publication editors
- Wanxiang Che
- Joyce Nabende
- Ekaterina Shutova
- Mohammad Taher Pilehvar
Abstract
Automatic language identification is a core problem of many Natural LanguageProcessing (NLP) pipelines. A wide variety of architectures and benchmarks havebeen proposed with often near-perfect performance. Although previousstudies have focused on certain challenging setups (i.e. cross-domain, shortinputs), a systematic comparison is missing. We propose a benchmark that allows us to test for the effect of input size, training data size, domain, number oflanguages, scripts, and language families on performance. We evaluatefive popular models on this benchmark and identify which open challengesremain for this task as well as which architectures achieve robust performance. Wefind that cross-domain setups are the most challenging (although arguably mostrelevant), and that number of languages, variety in scripts, and variety inlanguage families have only a small impact on performance. We also contributepractical takeaways: training with 1,000 instances per language and a maximuminput length of 100 characters is enough for robust language identification.Based on our findings, we train an accurate (94.41{\%}) multi-domain languageidentification model on 2,034 languages, for which we also provide an analysisof the remaining errors.
Publication metrics
PlumX, opens in new tab
Captures
6
Access to documents
Related Event
Title
Annual Meeting of the Association for Computational Linguistics
Event type
ConferenceDegree of recognition
International eventDate
27/07/2025 - 01/08/2025Location
ViennaAustria
