A test run of training
Echo-TTS on
DACVAE
Model doesn't generate coherent speech yet -- this is very much a test and don't expect to use it for anything (just use base Echo)
Todo:
- Convert model to Safetensors (get rid of dangerous file warning)
- Larger, multi-speaker dataset
- Scale up training to >1B parameters
Thanks to Hugging Face for making this run possible!