A Verifiable, Rule-Based Dataset for Reinforcement Learning with Verifiable Rewards (RLVR)
π Dataset Card π Usage βοΈ License
This dataset is designed to enhance the Instruction Following capabilities of Large Language Models (LLMs) through Reinforcement Learning (RL). Unlike subjective preference datasets (e.g., standard RLHF), this dataset focuses on Objective, Rule-Based Constraints.
Each entry provides a prompt with⦠See the full description on the dataset page:
https://huggingface.co/datasets/renhuimin/RL-Instruction-Following-Dataset.