This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.