LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation with Spoken Language Models" (arXiv 2024).
Clone the repo
git clone
git@hf.co:datasets/ilyakam/librispeech-long && cd librispeech-long
Process the dataset:… See the full description on the dataset page:
https://huggingface.co/datasets/ilyakam/librispeech-long.