German PronunCheck Mega Dataset 🇩🇪
Dataset Summary
This is a highly curated, 123GB+ mega-dataset designed specifically for training and fine-tuning German Automatic Speech Recognition (ASR) and Computer-Assisted Pronunciation Training (CAPT) models, such as HuBERT and Wav2Vec2.
Composition
This dataset is a clean concatenation of three distinct open-source datasets: