Do Syntactic Categories Help in Developmentally Motivated Curriculum Learning for Language Models?
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 288-300 (13 pages)Publication milestones
- Published - 01/11/2025
Publication status
Published - 01/11/2025
Place of publication
Suzhou, ChinaPublisher
Association for Computational Linguistics, United StatesISBN (Print)
TODOHost publication title
Proceedings of the First BabyLM WorkshopHost publication editors
- Lucas Charpentier
- Leshem Choshen
- Ryan Cotterell
- Mustafa Omer Gul
- Michael Y. Hu
- Jing Liu
- Jaap Jumelet
- Tal Linzen
- Aaron Mueller
- Candace Ross
- Raj Sanjay Shah
- Alex Warstadt
- Ethan Gotlieb Wilcox
- Adina Williams
Abstract
We examine the syntactic properties of BabyLM corpus, and age-groups within CHILDES. While we find that CHILDES does not exhibit strong syntactic differentiation by age, we show that the syntactic knowledge about the training data can be helpful in interpreting model performance on linguistic tasks. For curriculum learning, we explore developmental and several alternative cognitively inspired curriculum approaches. We find that some curricula help with reading tasks, but the main performance improvement come from using the subset of syntactically categorizable data, rather than the full noisy corpus.
Publication metrics
PlumX, opens in new tab
Captures
3
Access to documents
Related Event
Title
Workshop on BabyLM: Accelerating Language Modeling Research with Cognitively Plausible Data
Event type
ConferenceDegree of recognition
International eventDate
05/11/2025 - 09/11/2025Location
SuzhouChina
