text-to-speech model based on spark-tts, it supports English, Indonesian, Malay, Thai, Spanish, Tagalog
for inference, you can just ues the code from
https://github.com/SparkAudio/Spark-TTS ,just repalce the LLM model folder with this project.
inference with text prompt may cause some empty audio, can inference without text prompt, this can avoid the issues, but it may come at the cost of reduced performance.