Bilingual-BabyLM dataset collected as part of the paper "Bringing Up a Bilingual BabyLM: Investigating Multilingual
Language Acquisition Using Small-Scale Models" by Linda Zeng, Steven Y. Feng, and Michael C. Frank. For more details, please see Appendices A-B in our paper.
The dataset is a synthetic bilingual and code-switching dialogue dataset designed for training and evaluating small language models in multilingual and code-switched settings. It… See the full description on the dataset page:
https://huggingface.co/datasets/lindazeng979/bilingual-babyLM.