This dataset is outdated, its new version can be found here
Dataset Summary
This dataset was created with the OLaPh framework and used to train the OLaPh grapheme-to-phoneme model. It contains multilingual text from the FineWeb datasets, automatically phonemized into text–phoneme pairs for English, German, French, and Spanish.