APPS Dataset for Reinforcement Learning with AI Feedback
Dataset Details
APPS_RLAIF is an extended work from APPS [1]
to use Chat LLMs to create multiple variances for each solution for defined problems.
In each solution, we use LLama 34B [2] to transform the original solutions into variances and rank them by score.
The generated flow is demonstrated as below; each variance is created based on the previous version of it in the chat.
We iterated each solutions n=3 times… See the full description on the dataset page: https://huggingface.co/datasets/nmd2k/apps_rlaif.