This indonesian finetune of F5-TTS is made to introduce indonesian speech capabilities on the model.
Length: 43.35 hours Audio samples: 43999
Dataset sources: • data-indsp-news-lvcsr
The model has some difficulties in accurately matching the zero shot voice and emotions. The model also hallucinates on long texts.
Reference text: "Tidak ada yang menakutiku, bahkan kematian sekalipun."
Generated text: "Halo. Model… See the full description on the dataset page:
https://huggingface.co/datasets/yarengomik/new_dataset.