This dataset combines prompt-only datasets by capability theme for distillation experiments.
It contains 1,964,794 unique prompts from 5,827,983 raw rows;
3,863,189 exact canonical duplicates were removed.
Rows retain the canonical prompt-extraction columns and add source_repo_id for provenance.
Deduplication uses normalized system_prompt, prompt, tools, and schema_str, with the first
row in manifest order retained. Original source… See the full description on the dataset page:
https://huggingface.co/datasets/jamesdborin/Nemotron-Coding-and-Software-Engineering-prompt-only.