Subject-policy training corpus for the ph1 policy-horizon project: google/gemma-3-12b-it
greedy (T=0, 256 tok) responses to 40k deduped English WildChat first-turn prompts
(<=300 chars, template-deduped against the frozen eval set
cds-jb/ph1-policy-forks).
~5k rows additionally carry a two-turn continuation: gemma's reply after a mild pushback
(Are you sure? variants) appended to its own first answer… See the full description on the dataset page:
https://huggingface.co/datasets/cds-jb/ph1-gemma3-rollouts.