Complete rollout + reward record for the iter-2 GRPO run: every trajectory the policy
generated during training, with its graded reward. Preserved so the run stays re-analysable
after its torch_dist checkpoints were retired (only latest-3 survive; the 11 HF milestones
at iter_0/4/9/.../44/49 are the durable checkpoint record).
init
iter_49 of the iter-1 RL run… See the full description on the dataset page:
https://huggingface.co/datasets/sweagent/iter2-rl-rollouts.