Stage 2 Mixed Text/Phoneme TTS Dataset
This dataset contains mixed text/phoneme sequences for TTS training with curriculum learning.
The probability of converting words to phonemes increases over the dataset:
Start: p = 0.3 (more text, less phonemes)
End: p = 1.0 (all phonemes)
Transition: Linear over 500,000 rows
Each row uses p(i) for ALL its words/spaces, then i increments for the next row.
Column
Description