climbmix-split reorganizes the detokenized NVIDIA ClimbMix source corpus into
source-oriented splits. ClimbMix is described as being built from Nemotron-CC
and SmolLM-Corpus. Because SmolLM-Corpus is the smaller and directly
identifiable component, we used exact normalized-text matching against
SmolLM-Corpus to recover the SmolLM-derived portions. The remaining rows are
provided as the residual nemotron-cc split.
The data rows are unchanged from the detokenized… See the full description on the dataset page:
https://huggingface.co/datasets/AIML-TUDA/ClimbMix-split.