This arm uses solution context from cross_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training uses… See the full description on the dataset page:
https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution-xprob.