A unified Kazakh speech dataset for training text-to-speech (TTS) models. Compiled from two open-source ISSAI datasets — KazakhTTS and KazEmoTTS — with consistent preprocessing, quality filtering, and a single standardized schema.
Dataset Summary
Property
Value
Samples
232,350
Total audio
438.8 hours
Speakers
8 (5 professional studio + 3 emotional)
Emotions
6 (neutral, happy, sad, angry, scared, surprise)