A verifier-carrying 1:1:1 subsample of NVIDIA's open reasoning corpora,
built for Dr.GRPO / RLVR runs — 8000 train and 500 validation prompts
per domain.
code
nvidia/OpenCodeReasoning (split_0)
stdio_tests
program run on the… See the full description on the dataset page:
https://huggingface.co/datasets/Ksgk-fy/open-reasoning-rlvr-24k.