A faithful re-publication of the official Expresso
dataset (Nguyen et al., Interspeech 2023) as a loadable HuggingFace audio dataset, sourced
directly from FAIR's official tar.
⚠️ License: CC-BY-NC-4.0 — non-commercial use only.
read — 11.6k mono read-speech utterances with human transcripts.
conversational — ~15.9k mono per-utterance turns derived from the stereo conversational dialogues, transcribed with Whisper Large V3 Turbo.… See the full description on the dataset page:
https://huggingface.co/datasets/shangeth/expresso.