This project fine-tunes a TTS (Text-to-Speech) model using an mp3 file extracted from a YouTube video. The training was conducted on a Hugging Face Space running locally via Docker. A GPU is recommended for faster training.
Fine tuned with this docker image
FineTune Xtts Docker image
This model is based on xtts v2 which cannot be used commercially as per the
xtts license which is in a limbo state