Bulgarian BabyLM Dataset curated by Mila Marcheva (University of Cambridge).
28,467,275 tokens (excluding punctuation)
A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count.
We additionally release information about the source of sentence in the Bulgarian BabyLM Dataset, which you can find here:… See the full description on the dataset page:
https://huggingface.co/datasets/climb-mao/Bulgarian-BabyLM.