A unified, training-ready SFT dataset that assembles the DeepSeek-generated portions of two
public NVIDIA Nemotron datasets and normalizes them into a single OpenAI-style message schema
that renders directly with the DeepSeek-V4 chat encoding.
8,961,884 samples, 9 dataset×domain partitions, ~113 GB (zstd parquet, ~1 GB shards).
Covers math (CoT, tool-integrated reasoning, proofs), software-engineering, and
terminal-agent tasks.… See the full description on the dataset page:
https://huggingface.co/datasets/ycchen/nemotron-deepseek-sft-mix.