Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.
The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page:
https://huggingface.co/datasets/txchmechanicus/Nemotron-Terminal-Corpus.