This dataset contains 10,000 BoardGameBench prompt/reward examples generated for GRPO-style reinforcement learning on board-game move selection.
The final nemotron-boardgame-answer-lora-b4-safe-final adapter used this reviewed GRPO corpus after SFT and DPO. For that final pilot run, training used the first 512 examples from grpo_train.jsonl; the full 10k reviewed set is published here for reproducibility and follow-up training.
Format… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/boardgamebench-answer-grpo.