Based on the Hi-Fi Multi-Speaker English TTS Dataset, this dataset includes 148.06 hours and 75,575 utterances of 24kHz Mono 16-bit audio. It uses the utterances of speakers 6097, 6670, 6671, and 9017 from the original corpus.
For each original bona fide utterance, there are 4 corresponding deepfake zero-shot 1:1 utterances generated by the following models (15,115 utterances each):
F5-TTSv1
MaskGCT
CosyVoice2
Fish-S1mini
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grotsoV/4TTSDeepfakeData.