This repository contains 5M token subsets of the German BabyLM corpus that differ in their utterance-level construction distribution. For more information, please refer to the paper referenced below.
If you use this dataset, please cite the following publication:
@inproceedings{bunzeck-etal-2025-construction,
title = "Do Construction Distributions Shape Formal Language Learning In {G}erman {B}aby{LM}s?",
author = "Bunzeck, Bastian and
Duran, Daniel and
Zarrie{\ss}, Sina"… See the full description on the dataset page:
https://huggingface.co/datasets/bbunzeck/german-babylm-5m-subsets.