The Nemotron-Cascade-RM-Training dataset is designed for Reward Model (RM) training. It contains prompts and associated metadata to support the development of preference model for RLHF.
This dataset is ready for commercial use.
The dataset contains the following subset:
This data contains 81,808 samples used for RM training. It includes prompts, data sources, and category information.
This dataset is a curated subset of datasets… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RM-Training.