Training rollouts and evaluation outputs for the GRPO agent runs in this project.
Each rollout file is one optimizer step; each line is one sampled trajectory with
its decoded prompt, response and reward.
rollouts/search_r1_qwen3_8b_4gpu — Search-R1 / Qwen3-8B outcome-GRPO training rollouts
rollouts/search_r1_qwen3_8b_perturnnorm — Search-R1 / Qwen3-8B Process-GRPO (per-turn-norm) training rollouts
rollouts/alfworld_qwen3_8b_gigpo… See the full description on the dataset page:
https://huggingface.co/datasets/wckwan/PRM-agent-rl-artifacts.