This is a specially trained Coqui TTS Coqui TTS model specially for Sinhala, developed by Dialog Axiata PLC and the Dialog – UoM Research Lab.
We trained it on a custom recorded dataset adapting a strong male voice.
Features
Model architecture: VITS
Language: Sinhala (si-lk)
Training Sampling rate: 22050 Hz
Framework: Coqui TTS
Dataset
Voice: Male (Roshan)
Recording Sampling Rate: 44100Hz
No. of Clips: 1096
Total Length: >100mins (~2 hrs.)
Training Specs
Hardware: NVidia GeForce GTX1060 6GB GPU
Training Time: ~100 hours
Global Steps: 210,000
Batch Size: 16
Epochs:
Loss Convergence: Stable mel + KL losses
Installation
You can run this model locally using the included Flask-based inference server. This server will automatically use CUDA if it's available on your system.