Jumelet, Jaap and Fourtassi, Abdellah and Haga, Akari and Bunzeck, Bastian and Shandilya, Bhargav and Galvan-Sosa, Diana and Ghifari Haznitrama, Faiz and Padovani, Francesca and Meyer, Francois and Hu, Hai and Etxaniz, Julen and Prevot, Laurent and He, Linyang and Grandury, María and Marcheva, Mila and Foroutan, Negar and Theodoropoulos, Nikitas and Sadeghi, Pouya and Song, Siyuan and Salhan, Suchir and Zhou, Susana and Paniv, Yurii and Zhang, Ziyin and Bisazza, Arianna and Warstadt, Alex and Choshen, Leshem (2026) BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data, Proceedings of 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 24-29 March 2026, Rabat, Morocco, 3297-3329, Association for Computational Linguistics.
|
Text
2026.eacl-long.152.pdf Download (765kB) |
Abstract
We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally plausible pretraining data aiming to cover the equivalent of 100M English words of content in each of 45 languages. We compile evaluation suites and train baseline models in each language. BabyBabelLM aims to facilitate multilingual pretraining and cognitive modeling.
| Item Type: | Conference paper |
|---|---|
| Subjects: | Computing methodologies > Artificial intelligence > Natural language processing |
| Date Deposited: | 03 Sep 2026 07:13 |
| Last Modified: | 03 Sep 2026 07:13 |
| URI: | https://pubs.cs.uct.ac.za/id/eprint/1794 |
Actions (login required)
![]() |
View Item |
