The training data for drooryck/babylm-macaroni,
the BabyLM 2026 Multilingual-track submission. All corpora are created from the official BabyLM 2026
multilingual corpora (babylm-{eng,nld,zho}) and embedded with code-switching using an instruction-tuned LLM.
shuffled_cs/
Full code-switched corpus, all documents globally shuffled.
shuffled_nocs/
Matched unilingual twin (same documents… See the full description on the dataset page:
https://huggingface.co/datasets/drooryck/babylm-macaroni-corpus.