This dataset contains ~22,518 training triplets for fine-tuning a zero-shot voice+emotion cloning TTS model. Each sample provides everything needed to train a model that can clone both a speaker's voice identity AND their emotional delivery from separate reference audio clips.
The data is stored as WebDataset .tar shards, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-emo-cloning-dataset.