A huge dataset of duplicated code issue pairs from Github. Over 1.2M issues, composing over 750K primary-duplicate issue pairs.
This dataset contains pairs of code issues raised in open-source repositories on Github, where one issue was identified to be the duplicate of another.
For both the primary issue and the duplicate, the title and body are provided. The dataset spans 25,000 of the most popular open-source repositories, containing over 1.2M… See the full description on the dataset page:
https://huggingface.co/datasets/primordic/IssueDupes.