Diagnosing Component-Level Failures in Computer-Use Agents
ComponentBench is a diagnostic benchmark for computer-use agents that targets the middle layer between atomic GUI-grounding tests (e.g., ScreenSpot) and long-horizon workflow benchmarks (e.g., WebArena, OSWorld). It evaluates agents on individual UI component interactions — toggling button groups, setting sliders, using date pickers — that are short enough to diagnose specific failures but rich enough to… See the full description on the dataset page:
https://huggingface.co/datasets/TianchenGuan/ComponentBench.