Data accompanying the anonymous submission, "Generative Verifiers: Reward Modeling as Next-Token Prediction".
Includes the verification rationales used to finetune the generative verifiers.
This is a public huggingface repo of the dataset available on Github. The original paper is here with an accompanying website here
Zhang et al. (2025). Generative Verifiers: Reward Modeling as Next-Token Prediction.
GSM8K data: Copyright (c) 2021 OpenAI… See the full description on the dataset page:
https://huggingface.co/datasets/flowingpurplecrane/genrm.