Pilot corpus of unlabeled speech: audio segmented by energy-based voice
activity detection (VAD), with speaker metadata. No transcriptions are provided at this stage.
Intended for self-supervised pre-training (wav2vec 2.0, MMS, HuBERT) and as a
basis for annotation batches.
This corpus is the initial pilot of a wider programme building speech
resources for low-resource languages. It is released to document the method —
segmentation, quality… See the full description on the dataset page:
https://huggingface.co/datasets/labari-voice/fon-speech-pilot.