OleSpeech-IV-2025-EN-AR-100 is a publicly released subset of OleSpeech-IV, a large-scale multilingual and multispeaker conversational speech dataset. This release is intended for non-commercial research use and provides high-quality speech data aligned with transcripts, speaker turns, and additional metadata. A sample of the dataset can be found in the sample/ folder.
OleSpeech is a tiered dataset series designed to provide speech… See the full description on the dataset page:
https://huggingface.co/datasets/olewave/OleSpeech-IV-2025-EN-AR-100.