Real-world white-collar tasks across ~35 professions. Each task has a prompt,
reference files, a weighted rubric, and a human-readable brief.
The dataset ships two splits:
main — 65 full tasks. Some include a files_required_to_search/ folder
of ground-truth materials the agent is expected to discover via search.
easy — 63 simplified tasks (shorter prompts, no files_required_to_search/).
Useful for cheaper smoke tests and capability ranking.
The main and easy splits… See the full description on the dataset page:
https://huggingface.co/datasets/JobBench/job-bench.