This dataset is used to evaluate Multi-Line fill in the middle code completion capabilities of a system.
The dataset is derived from SWE-Bench dataset.
Evaluation is performed by stiching the generated middle portion, with the other patch and passing into the SWE Evaluation harness, which runs unit test verification and calculate Pass@1.
In addition to the fields already calculated by SWE-Bench dataset, this dataset contains five… See the full description on the dataset page:
https://huggingface.co/datasets/sourcegraph/code-multi-line-infilling-benchmark.