Abstract
We present Kenji-Endo, a BabyLM pretrained on a dedicated Italian dataset, which participated in four tasks at the 9th edition of EVALITA: DeSegMa, MultiPRIDE, IMPOLS, and FadeIT. Kenji-Endo achieved competitive performance across all tasks, demonstrating that language modeling with limited data and compact model sizes
can represent a viable alternative to Large Language Models
can represent a viable alternative to Large Language Models
| Original language | Undefined/Unknown |
|---|---|
| Title of host publication | Ceur Workshop Proceedings |
| Number of pages | 11 |
| Publisher | CEUR Workshop Proceedings |
| Publication date | 2026 |
| Pages | 16-26 |
| Publication status | Published - 2026 |
| Externally published | Yes |
| Event | 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop - Bari, Italy Duration: 26 Feb 2026 → 27 Feb 2026 Conference number: 9 https://www.evalita.it/campaigns/evalita-2026/ |
Conference
| Conference | 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop |
|---|---|
| Number | 9 |
| Country/Territory | Italy |
| City | Bari |
| Period | 26/02/2026 → 27/02/2026 |
| Internet address |
Keywords
- NLP
- Evaluation
- Italian
- Language Models
- BabyLM
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver