370 tasks representing realistic developer work in large, enterprise-scale
codebases, spanning 20 developer work types. Each row is the
problem layer of one task: the instruction an agent receives, a pointer to the
repository snapshot and base commit, and grading metadata. The repository
contents and verifier environment are materialized at evaluation time by the
CodeScaleBench eval kit, not shipped in this dataset.
This is the frozen csb-v1-mixed371 analysis set… See the full description on the dataset page:
https://huggingface.co/datasets/sgjarmak/CodeScaleBench.