IHBench evaluates post-interruption recovery in voice agents executing
structured, multi-step workflows. Unlike benchmarks that measure the timing of
interruptions (barge-in detection, endpointing, turn-taking), IHBench measures
what the agent says after an interruption: does it resume the workflow at the
correct step, address the user's interjection, and avoid re-delivering content the
user already heard?
The benchmark contains 45… See the full description on the dataset page:
https://huggingface.co/datasets/bosonai/IHBench.