Note: This dataset has an improved and significantly larger successor: SWE-rebench.
SWE-bench Extra is a dataset that can be used to train or evaluate agentic systems specializing in resolving GitHub issues. It is based on the methodology used to build SWE-bench benchmark and includes 6,415 Issue-Pull Request pairs sourced from 1,988 Python repositories.
The SWE-bench Extra dataset supports the development of software engineering agents… See the full description on the dataset page:
https://huggingface.co/datasets/nebius/SWE-bench-extra.