This is the official repository for the paper "FeedbackEval: Evaluating Large Language Models in Feedback-Driven Code
Repair".
We construct a new benchmark, FeedbackEval, to systematically evaluate LLMs’ ability to interpret and
utilize various feedback types in code repair.
FeedbackEval consists of 394 coding tasks covering a diverse range of programming scenarios. In total… See the full description on the dataset page:
https://huggingface.co/datasets/SYSUSELab/FeedbackEval.