Three disjoint random subsets of 10,000 English clips each (30,000 total,
~77.5 h) sampled from amphion/Emilia-Dataset
(Emilia/EN/*.tar), for use as inference-time TTS (ITTS) generation prompts.
Each row includes the original Emilia audio (mp3, 24 kHz) plus its transcript
and metadata.
from datasets import load_dataset
ds =… See the full description on the dataset page:
https://huggingface.co/datasets/dlion168/emilia-en-itts-prompts.