DeepResearch Benchmark 2.0 is a collection of 100 English deep-research benchmark cases.
Each case asks a model to analyze 6-10 entities across 6-10 research dimensions, and includes:
the public user-facing question,
a reference answer with derivations and source URLs,
a detailed scoring rubric,
metadata for the generation/auditing pipeline when available.
This Hugging Face package is the clean OpenReview dataset release. It excludes local MCP configs… See the full description on the dataset page:
https://huggingface.co/datasets/xsx001/deepresearch-benchmark-2.