🎙️ Bengali-Loop: A Long-Form Bangla Speech Corpus
Dataset Summary
Bengali-Loop is a large-vocabulary, long-form Bangla (Bengali) speech corpus designed to push the boundaries of Automatic Speech Recognition (ASR) in low-to-mid resource settings. It comprises 155 hours of naturally occurring Bangla speech sourced from 249 YouTube videos spanning drama serials, audiobooks, and entertainment channels — making it one of the most diverse publicly available Bangla ASR datasets… See the full description on the dataset page: https://huggingface.co/datasets/Suprio85/Bangla_Speech_Corpus.