A dataset of 10,000+ AI-generated (deepfake) speech samples created for training and evaluating deepfake speech detection models.
Source subsets
clean (train.100, train.360, validation, test) + other (train.500, validation, test)… See the full description on the dataset page:
https://huggingface.co/datasets/gereon/voxguard-synthetic-speech.