A large-scale benchmark built on 47,214 research papers from five top machine learning conferences, including full text, peer reviews, editorial decisions, GROBID-parsed metadata, and bibliographic references.
PRISM evaluates LLM-based peer reviewers across five dimensions — Validity, Helpfulness, Comprehensiveness, Specificity, and Faithfulness — and supports tasks including review generation, meta-review… See the full description on the dataset page:
https://huggingface.co/datasets/anoyresearcher/prism_paper_data.