Purpose: Docker replay verdicts (with independently-derived evidence,
not the source's own self-reported task_complete flag) over
nvidia/Nemotron-Terminal-Corpus's skill_based configs, for small-model
SFT on CLI/terminal-agent behavior. Only VERIFIED trajectories are usable
as direct SFT material; the rest are documented failure/gap evidence.
Upstream: nvidia/Nemotron-Terminal-Corpus
@ a1667c4ffdadea02a89bffe4f1bb7ca2ff19f8d9, configs… See the full description on the dataset page:
https://huggingface.co/datasets/EinicherAI/nemotron-terminal-replay-evidence.