A more aggressively cleaned up version of Calvin-Xu/Furigana-Aozora-Speech, which consists of 2,536,041 out of the 3,361,443 entries generated from the raw data 青空文庫及びサピエの音声デイジーデータから作成した振り仮名注釈付き音声コーパスのデータセット
https://github.com/ndl-lab/hurigana-speech-corpus-aozora. Training data of Calvin-Xu/FLFL.
The Whisper-generated transcriptions in the original dataset contains many errors. Additional sanity-checking is implemented to filter on reading of common kanji and eliminate wildly inaccurate… See the full description on the dataset page:
https://huggingface.co/datasets/Calvin-Xu/FLFL-Aozora-Speech-Train.