This dataset is part of a Korean multispeaker speech corpus project for Text-to-Speech (TTS) research.It includes preprocessed speech–text pairs, metadata, and linguistic annotations for model training.
This dataset is designed as a robustness test set for evaluating Korean TTS models.It is divided into three subsets:
Clean – clean and high-quality speech recordings
Numeric – speech containing a high proportion of numbers
Noisy – speech with environmental… See the full description on the dataset page:
https://huggingface.co/datasets/aanonyyy/M6A5P1Q7.