Complete training artifacts for the pre-registered reward shaping ×
opponent diversity factorial study in
AI-Plays-Tag: 195 PPO
zoo-self-play training runs (plus three 10M-step pilots and the 4
pre-registered anchor runs), 5M steps each, in a 2D tag environment.
Headline result (confirmatory, pre-registered): reward shaping and zoo
opponent mixing are complements — β_RA = +2.05 log-odds
[+0.72, +3.45], P(>0) = 0.999 — and dense shaping trades… See the full description on the dataset page:
https://huggingface.co/datasets/kilojoules/ai-plays-tag-design-c.