Built by build_tts_dataset.py (companion of the ASR mix builder). One
config per source; every config has the same schema (audio at its
NATIVE sampling rate — see sampling_rate — plus text, language,
speaker_id, source, license, domain, duration_s).
Audio bytes are copied from the sources unmodified, so sample rates are
mixed across configs (see the table). Never blanket-resample the union
upward — upsampled 16 kHz audio has no… See the full description on the dataset page:
https://huggingface.co/datasets/PrinceAlhassanNasamu/tekyerema-pa-tts.