A multi-task code reinforcement-learning dataset mixture in the
verl RL prompt format. It pairs a
code-generation split with a suite of auxiliary code-understanding tasks so the
same corpus can drive three training regimes from one repo. It is the
V3-dedupe successor to OctoReasoner/FinalMix (see
Relationship to FinalMix (v1)).
train_aux_cascade
25… See the full description on the dataset page:
https://huggingface.co/datasets/OctoReasoner/FinalMix2.