Full SFT dataset for terminal-agent fine-tuning, derived from
nvidia/Nemotron-Terminal-Corpus
(config skill_based_medium), converted to the terminus-2 "thinking-preservation"
chat format and reproducibly shuffled. 89,343 multi-turn agent trajectories across
11 terminal skills, ready to train with the AReaL
SFT recipe in ethanewer/posttraining-2606.
This is the dataset used by config_terminus2_l40s_default.yaml in that repo. Pair it
with the base… See the full description on the dataset page:
https://huggingface.co/datasets/eewer/skill-based-medium-terminus2-sft.