A privileged variant of PRIME-RL/Eurus-2-RL-Data,
built for privileged-information RL experiments (giving the value network / process reward
model access to a worked reference solution z that the policy never sees).
The train split adds two columns; every other column is preserved verbatim from the
base Eurus-2-RL-Data:
reference_solution
the full worked reference… See the full description on the dataset page:
https://huggingface.co/datasets/jasonkena/Eurus-2-RL-Data-privileged.