Nemotron Nano RL Math 22K is a 22,056-example English mathematical reasoning dataset prepared for reinforcement learning with verifiable rewards (RLVR). It contains the DAPO-Math and Skywork math profiling components identified in the metadata of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented in a compact prompt / label / metadata JSONL schema.
Each example contains a ready-to-use user prompt, a reference answer, and empirical pass-rate… See the full description on the dataset page:
https://huggingface.co/datasets/wflying/nemotron-nano-rl-math-22k.