For usage instructions see the
F5-TTS repo and use
F5-TTS Base as base model.
Sample audios for voice reference should be about 6-12 seconds long. Longer clips should be clipped to avoid automatic clipping in the middle of a word.
If you get errors in generated audio check thet the automatic transcript of the reference audio is correct and adjust if needed.
Model is trained on Mozilla Common Voice 24.0 dataset.