GitGoodBench Lite is a subset of 900 samples for evaluating the performance of AI agents in resolving git tasks (see Supported Scenarios).
The samples in the dataset are evenly split across the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit gram.
This dataset thus contains 150 samples per sample type and programming language.
All data in this dataset are collected from 479 unique, open-source GitHub… See the full description on the dataset page:
https://huggingface.co/datasets/JetBrains/git_good_bench.