A benchmark for evaluating code review agents on real-world GitHub issues with executable verification. Each instance pairs a GitHub issue with an AI-generated pull request; the reviewer must decide whether the PR resolves the issue and, if not, provide structured feedback to guide revision.
Project Page | Paper | Code | Benchmark | Training Data | Claude Code Plugin
SWE-Review is a framework for closing the issue-resolution… See the full description on the dataset page:
https://huggingface.co/datasets/Lego-X/SWE-Review-Bench.