A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and French BabyBabelLM corpora.
Sampling is stratified by category: each source category is sampled independently to preserve the original category proportions within each language.
Token counts
Language
Tokens
Share
English (eng)
49,481,353
43.9%
French (fra)
63,337,862
56.1%
Total112,819,215
100%
Category… See the full description on the dataset page: https://huggingface.co/datasets/adzcai/babylm-eng-fra-50-50-stratified.