We collected 7,777 successful multi-turn trajectories by running
GLM-5.1 with the
PI agent on
nebius/SWE-rebench-V2.
Each SWE-rebench-V2 task is an issue of a public repository. An agent has to
implement a fix that is then validated by unit tests.
The model was run on 3,837 issues from Python repositories, with 4 rollouts
per issue. The task list is available as the train split of… See the full description on the dataset page:
https://huggingface.co/datasets/whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces.