Arxiv Paper | Hugging Face Dataset | Evaluation Code
We propose a formal characterization of the deep research (DR) problem and introduce a new benchmark, LiveDRBench, to evaluate the performance of DR systems. To enable objective evaluation, we define DR using an intermediate output representation that encodes key claims uncovered during search—separating the reasoning challenge from surface-level report generation.… See the full description on the dataset page:
https://huggingface.co/datasets/microsoft/LiveDRBench.