ⓍTTS is a Voice generation model that lets you clone voices into different languages by using just a quick 6-second audio clip. There is no need for an excessive amount of training data that spans countless hours.
This is the same or similar model to what powers Coqui Studio and Coqui API.
Supports 17 languages.
Voice cloning with just a 6-second audio clip.
Emotion and style transfer by cloning.
Cross-language voice cloning.
Multi-lingual speech… See the full description on the dataset page:
https://huggingface.co/datasets/kingadil/super_tts.