Multi-turn terminal-agent conversations distilled from
Zhongzhi1228/Recursive-Task-Synthesis-Trajectories,
ready for supervised fine-tuning of Qwen/Qwen3.5-27B.
Pipeline, launchers, and the full plan:
https://github.com/k1ssloo/RST-Train
The source release has 327,189 trajectories. cap10 ends at 10,778 examples —
the count arXiv:2608.05466v3 states it trained
on. That was… See the full description on the dataset page:
https://huggingface.co/datasets/NiuNiu0110/RST-SFT-Qwen3.5-27B.