DuMateBench is a benchmark dataset for evaluating AI agents on realistic
computer-based work tasks.
Each task provides an instruction, a sandboxed workspace, task-specific
resources, and an evaluator. The agent must inspect the workspace, use
available tools, produce the required artifact, and recover from environmental
or tool failures when necessary.
The dataset contains 200 tasks covering software development, web research… See the full description on the dataset page: https://huggingface.co/datasets/Annihi/dumate_bench.