This is a preprocessed dataset made by applying Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing (
https://huggingface.co/spaces/sarulab-speech/sidon_demo_beta) on LJSpeech-1.1 dataset
Format: Following Stylish-TTS
Files:
Reproduction code:
%cd /content
!sudo apt install aria2 -y
!rm LJSpeech-1.1.tar.bz2
!aria2c -x 16
https://data.keithito.com/data/speech/LJSpeech-1.1.tar.bz2
!tar -xf… See the full description on the dataset page:
https://huggingface.co/datasets/hr16/ljspeech-sidon.