SWE-PRBench is a benchmark of 350 pull requests with human-annotated
ground truth for evaluating whether LLMs can identify the same issues
that real human reviewers flag in production code.
Existing benchmarks like SWE-Bench measure whether models can produce
correct code. SWE-PRBench… See the full description on the dataset page:
https://huggingface.co/datasets/foundry-ai/swe-prbench.