The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - Get the data
Description
An extended benchmark of real-world software engineering tasks with enhanced… See the full description on the dataset page:
https://huggingface.co/datasets/ud-nlp/swe-bench-coding-tasks.