This dataset card shows details about the Yapdo conversational speech corpus by Liva AI (YC S25). This dataset card details information for 30,000+ hours of recordings from 8,000+ speakers across 17 languages, with the rest of the hours still undergoing QA (estimated 50k total). The source audio is natively recorded with separate speaker channels; the samples here are presented as combined conversations.
The strength of this dataset is its naturalness. Recorded… See the full description on the dataset page:
https://huggingface.co/datasets/liva-ai/yapdo-convo.