SFT dataset used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
This dataset is used for CapTTS, EmoCapTTS and AccCapTTS tasks.
Please refer to CapSpeech for the whole dataset.
audio_path
string
File path to the audio sample. The actual audio is hosted separately.
text
string
The transcript corresponding to the audio sample.
source
string
The original dataset… See the full description on the dataset page:
https://huggingface.co/datasets/OpenSound/CapTTS-SFT.