Hypa-LibreSpeech is a large-scale, multilingual speech dataset curated by Hypa AI for training and evaluating Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems across 8 European languages. It contains 200,000 high-quality audio–text pairs derived from open-domain audiobook recordings originally sourced from the LibriVox project.
This dataset builds directly upon two foundational open-source corpora:
openslr/librispeech_asr — the… See the full description on the dataset page:
https://huggingface.co/datasets/hypaai/Hypa-LibreSpeech.