The task datasets for
WorkBuddy Bench, a benchmark
for evaluating coding agents on real-world developer, PM, algo, QA, ops, and
security work. This repository hosts only the task data; the evaluation
framework, setup, and usage all live in the GitHub repository above.
WorkBuddy Bench is built from real-world work tasks. Given a task and a
sandboxed workspace, an agent must produce the correct change — a patch, an
artifact… See the full description on the dataset page:
https://huggingface.co/datasets/tencent/workbuddy-bench.