This is a single-speaker synthetic speech dataset for American English. It carries the Dii voice, a female voice. The audio was synthesized with text-to-speech and adapted to the Dii speaker identity, to train a Dii voice model for American English. The dataset contains 1150 recordings, with a metadata.csv transcript file and one WAV file per line.