Views
No views yet
geodesic-research/nemotron-super-120b-cc-mt-uniform-nodr-200k-sft
RL-trained on GSM8k, to measure whether math RL post-training erodes the
alignment interventions applied during midtraining.openai/gsm8k train split (7,473 problems), verifiable
reward via HF math_verify (boxed/numeric answer match, judge-free),
boxed-answer CoT prompt (examples/prompts/cot.txt), no think prefill.env/accuracy, train/reward, pass@8, sampled completions).core/configs/isambard/gsm8k/ in geodesic-nemo-rl
(PR #67). Intermediate adapters every 50 steps are retained on-cluster.