This is a test dataset compatible with the ReinforceNow Platform containing only 92 sample entries for demonstration and testing purposes. It is designed to help users get started with RLHF training for mathematical reasoning models.
Note: This is not a production dataset. It contains a small subset of math problems for testing the ReinforceNow CLI and verifying your training pipeline works correctly before scaling up.
Data… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/rl-single-math-reasoning.