The YouTube Corpus of Singapore English Podcasts (YCSEP) contains ASR transcripts and audio from 620 hours of over 1,300 podcast episodes by Singapore-based content creators, comprising 756k individual turns and 8.38 million word tokens.
YCSEP was created using a pipeline comprising yt-dlp, WhisperX, and pyannote.audio, and is intended to advance the study of the linguistic and discourse… See the full description on the dataset page:
https://huggingface.co/datasets/stcoats/YCSEP_v1.