Single-speaker Bengali (bn) speech chunks extracted from YouTube videos,
cleaned and segmented for TTS (text-to-speech) training data.
Downloading source audio via yt-dlp
Trimming the first/last 2 minutes of each source video
Removing background music/noise via Demucs vocal separation
Splitting stereo channels into independent left/right tracks when present
Voice Activity Detection (Silero… See the full description on the dataset page:
https://huggingface.co/datasets/MeghanaKap/scraped_datasets.