To evaluate our model's performance, we have constructed a comprehensive evaluation dataset covering diverse practical scenarios. This dataset combines:
Tasks from CC-Bench: We selected tasks where we could match input source files to the projects and where the queries were clear and well-defined.
Internal Testing Data: We supplemented the dataset with additional tasks to increase data diversity, primarily including complex project… See the full description on the dataset page: https://huggingface.co/datasets/nex-agi/coding-eval.