This dataset is a clean GRPO/RLVR prompt set for the three weak Nemotron
challenge types identified after the 0.84 SDPO adapter diagnostics:
bit_manipulation, unit_conversion, and gravity.
The training rows are intentionally modeled as:
prompt x + gold answer r + verifier/reward spec
There are no source CoT traces, teacher completions, SDPO samples, RLSD
privileged traces, or eval predictions in the training split. GRPO should
sample completions… See the full description on the dataset page:
https://huggingface.co/datasets/dvyomkesh/nemo-grpo-weak3-from084-prompts.