PaperBench is a benchmark dataset for evaluating the ability of AI agents to replicate state-of-the-art AI research from scratch. The dataset contains 20 ICML 2024 Spotlight and Oral papers, each decomposed into hierarchical rubrics with clear grading criteria.
20 research papers from ICML 2024
8,316 individually… See the full description on the dataset page:
https://huggingface.co/datasets/josancamon/paperbench.