EVA-Bench is an end-to-end evaluation framework for conversational voice agents that orchestrates bot-to-bot audio conversations and scores them on both task accuracy and interaction experience.
No existing benchmark jointly addresses the two core evaluation challenges for voice agents: generating realistic simulated conversations, and measuring quality across the full scope of… See the full description on the dataset page:
https://huggingface.co/datasets/ServiceNow-AI/eva-bench.