A detailed dataset description (including description, audio samples, and statistics) is provided here:
https://domklement.github.io/sbcsae/
If you use the dataset, please, do not forget to cite our work:
@inproceedings{maciejewski24_interspeech,
title = {Evaluating the Santa Barbara Corpus: Challenges of the Breadth of Conversational Spoken Language},
author = {Matthew Maciejewski and Dominik Klement and Ruizhe Huang and Matthew Wiesner and Sanjeev Khudanpur},
year = {2024}… See the full description on the dataset page:
https://huggingface.co/datasets/dklement/SBCSAE.