Agentic tool-use traces collected from a GLM-5.2 teacher solving Harvey LAB
legal benchmark tasks inside the open-rl scaffold (bash / read / write / todo tools,
sandboxed workspace, 163,840-token trajectory budget, 32k max tokens per turn).
v2: the teacher's chain-of-thought is captured per turn in the reasoning
field (--reasoning-parser on the serving endpoint), so students can be trained
to think before acting — SFT on the… See the full description on the dataset page:
https://huggingface.co/datasets/ShubyM/harvey-lab-glm-traces.