This repository contains the datasets presented in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the issue of "too-easy" prompts by automatically composing multiple verifiable problems into a single, more challenging yet still verifiable prompt. RL training on these compositional prompts helps… See the full description on the dataset page:
https://huggingface.co/datasets/xx18/Composition-RL-EVA.