RebuttalBench is derived from the Re2-rebuttal dataset (Zhang et al., 2025), a comprehensive corpus containing initial scientific papers, their corresponding peer reviews, and the authentic author responses. The raw data undergoes a multi-stage processing pipeline: first, GPT-4.1 parses all the reviews into over 200 K distinct comment-response pairs, achieving a 98 % accuracy as verified by manual sampling; next, each review and comment… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/RebuttalBench.