Run the script from the repo root to create the parquet and copy audio files into data/audio_files:
python3 scripts/create_parquet.py
data/dataset.parquet (contains columns: source, text, audio)
data/audio_files//.wav (copied audio files)
The audio column contains relative paths starting with data/audio_files/... so the whole data/ folder can be uploaded to Hugging Face or copied… See the full description on the dataset page:
https://huggingface.co/datasets/Aybee5/ha-tts-mixed.