Fine-tuned Spark-TTS model for high-quality Kazakh language text-to-speech synthesis with voice cloning capabilities.
This model is specifically optimized for Kazakh language TTS, supporting both Cyrillic and Tote Zhazu (Arabic) scripts. It provides natural-sounding speech synthesis and voice cloning with just 3-10 seconds of reference audio.
This project is built upon the
Spark-TTS framework and is distributed under the
CC BY-NC-SA 4.0 license.