SpeechEval is a large-scale multilingual dataset for general-purpose, interpretable speech quality evaluation, introduced in the paper:
It is designed to train and evaluate Speech LLMs acting as “judges” that can explain their decisions, compare samples, suggest improvements, and detect deepfakes.
Utterances: 32,207 unique speech clips
Annotations: 128… See the full description on the dataset page:
https://huggingface.co/datasets/yanyan666/SpeechEval.