RebuttalBench is a benchmark for scientific agents to automatically implement new experiments from paper repositories, sourced from accepted-paper rebuttals. Evaluation is automated with an agent-as-a-judge protocol: the judge checks whether produced experiment result tables fully support scientific claims distilled from ground-truth rebuttal experiments.
The current release focuses on papers from ICLR 2026 and NeurIPS 2025.
Two task… See the full description on the dataset page:
https://huggingface.co/datasets/Jiefuo/RebuttalCodeBench.