This is the corrected p4_dataset_v2 preference dataset built from the
validity-gated P2 negotiation campaign. Each row pairs the model's stored
per-turn rendered prompt and verbatim action with an executable
best-response-oracle action rendered in the same wire format.
All 12,761 rows contain the required experiment-name field, set to
p4-divergence-dpo-v2.
train
5,533
Episode-level training… See the full description on the dataset page:
https://huggingface.co/datasets/siddharthmb/2026.RA.Divergence-DPO-Pairs.