OmniAgentBench is a benchmark for evaluating multimodal agents under realistic "wild" conditions: speech input, acoustic noise, dense/scattered instructions, and multi-turn conversations. It wraps three existing agent benchmarks (MPCC, GUI Odyssey, EmbodiedBench) with speech audio, noise overlays, and wild text rewrites so that the same tasks can be evaluated under controlled input-modality variations.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/omniagentbenchspeech/OmniAgentBench.