Campaign artifacts for branch-backtracking GRPO (Qwen3-8B LoRA negotiation policy): mine decision nodes
from saved no-deal episodes, replay the exact prefix (interlens.arena.replay.apply_prefix), sample K=8
on-policy continuations per node, and train on within-node advantages. Proposal, code, and the results note
(0063) live in the source repo; the preregistered POSITIVE gate failed — at ~10%… See the full description on the dataset page:
https://huggingface.co/datasets/siddharthmb/2026.RA.Branch-Backtrack-GRPO.