π Submitted to the Uncharted Data Challenge
hosted by Adaption Labs β credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale synthetic multilingual speech dataset β 68,677 clips across
9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base
using zero-shot voice cloning from 5 reference speakers.
Intended for training and evaluating TTS, ASR, voice conversion, and
multilingual speech models. Each clip is paired with the⦠See the full description on the dataset page:
https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.