BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data

Jumelet, Jaap and Fourtassi, Abdellah and Haga, Akari and Bunzeck, Bastian and Shandilya, Bhargav and Galvan-Sosa, Diana and Ghifari Haznitrama, Faiz and Padovani, Francesca and Meyer, Francois and Hu, Hai and Etxaniz, Julen and Prevot, Laurent and He, Linyang and Grandury, María and Marcheva, Mila and Foroutan, Negar and Theodoropoulos, Nikitas and Sadeghi, Pouya and Song, Siyuan and Salhan, Suchir and Zhou, Susana and Paniv, Yurii and Zhang, Ziyin and Bisazza, Arianna and Warstadt, Alex and Choshen, Leshem (2026) BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data, Proceedings of 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 24-29 March 2026, Rabat, Morocco, 3297-3329, Association for Computational Linguistics.

[thumbnail of 2026.eacl-long.152.pdf] Text
2026.eacl-long.152.pdf

Download (765kB)

Abstract

We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally plausible pretraining data aiming to cover the equivalent of 100M English words of content in each of 45 languages. We compile evaluation suites and train baseline models in each language. BabyBabelLM aims to facilitate multilingual pretraining and cognitive modeling.

Item Type: Conference paper
Subjects: Computing methodologies > Artificial intelligence > Natural language processing
Date Deposited: 03 Sep 2026 07:13
Last Modified: 03 Sep 2026 07:13
URI: https://pubs.cs.uct.ac.za/id/eprint/1794

Actions (login required)

View Item View Item