This is motivated by AlpacaEval, but it is aimed at evaluating advanced speech generation capabilities of end-to-end audio LLMs.
Instructions are curated manually (with the assistant of ChatGPT)
Instruction Audio is obtained from kokoro TTS
Output Audio is obtained from GPT-4o-Audio, Gemini-2.0-Flash-exp, Moshi, Typhoon2-Audio, DiVA+TTS, Qwen2-Audio+TTS.
Each row consists of <instruction, audio_a, audio_b> -- all in audio format (wav)
label = a, b… See the full description on the dataset page:
https://huggingface.co/datasets/vivavault/speakbench-v1-label.