A sample package of a multilingual software engineering task benchmark dataset, designed to evaluate AI Agents' capabilities in code fixing and feature implementation on real open-source projects. Runs on the Harbor evaluation framework.
This dataset contains 46 tasks covering 9 programming languages and 6 task types, sourced from real open-source repository commits. Each task provides a Chinese problem_statement (problem… See the full description on the dataset page:
https://huggingface.co/datasets/PipelineLabs/Multi2lingual_SWEBENCH_DEMO.