The training set used for the Scale-SWE on-policy self-distillation (OPSD) runs. 3200 SWE tasks across
752 repositories, each paired with a reference agent trajectory and a condensed solution hint.
Uploaded from /checkpoint/huggingface/datasets/scaleswe_opsd_v2_3200_summary (a
datasets.save_to_disk directory), converted to parquet. Row count, ids and field contents verified
identical to the source.
⚠️ Contains… See the full description on the dataset page: https://huggingface.co/datasets/starli-snowflake/scaleswe-opsd-v2-3200-summary.