Views
No views yet
NVIDIA-Nemotron-3-Super-120B-A12B-Base-Chat-Init-BF16 (pure Base weights;
1,188 chat-scaffolding embedding rows grafted from Instruct. No post-training behavior included.)data/curriculum_noDR_500M.jsonl (~509M tokens, consumed in exact file order, no-deliberative-reasoning 500M variant) and shuffled pretraining replay (geodesic-research/Nemotron-Pretraining-Specialized)geodesic-research/nemotron-base-tokenizer (EOD=</s>=id 2). All data filtered against
the 1,188 zero-embedding token ids of Super-Base (0 docs dropped).NemotronHForCausalLM; tokenizer PreTrainedTokenizerFast.
Coherence-checked post-export (8/8 non-empty, fluent completions).nemotron-super-120b-cc-mt-curriculum-nodr-500m-sr-sft (reasoning) and
nemotron-super-120b-cc-mt-curriculum-nodr-500m-200k-sft (non-reasoning).