This dataset provides one of the most comprehensive audio deepfake detection resources available, consisting of real and synthetic audio samples standardized for machine learning research. It is designed to support development and benchmarking of deepfake detection algorithms.
Dataset Composition
Total audio files: 189,221
Real: 101,172 (53.5%)
Fake: 88,049 (46.5%)
Class ratio (Real:Fake): 1.15:1