DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page:
https://huggingface.co/datasets/OpenSound/CapSpeech_Emilia.