A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation.
Dataset Details
Dataset Description
The FEA-Bench is a benchmark with a test set that contains 1,401 task instances from 83 Github repositories. This benchmark aims to evaluate the capabilities of repository-level incremental code development. The task instances are collected from Github pull requests, which have the purpose of new feature… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FEA-Bench.