This dataset consists of manually aligned audio–text pairs extracted from Kazakh songs and designed for research in automatic speech recognition (ASR) for low-resource languages. The primary goal of the dataset is to investigate whether sung speech can serve as a complementary training resource for Kazakh ASR systems.
The corpus contains line-level vocal segments obtained from commercially released songs, with manually verified… See the full description on the dataset page: https://huggingface.co/datasets/yeshpanovrustem/kazakh_songs_asr.