This project studies how nine tested LLMs score fixed review evidence, with a complete
eight-model paper-reading panel. The current study uses three venue-years built from official OpenReview snapshots. It is
designed to support two bounded questions:
How much do models' scoring standards differ when they receive the same human evidence?
Where the official review form separates strengths and weaknesses, does itemizing that… See the full description on the dataset page:
https://huggingface.co/datasets/stanjsx/LLMPeerReview.