Frozen eval set for the ph1 policy-horizon project: prompts on which subject
S = google/gemma-3-12b-it and auditor bases Qwen/Qwen3-{0.6B, 1.7B, 4B, 8B, 14B}
produce semantically divergent continuations under greedy decoding (T=0, 256 tokens),
mined from ~8.4k curated candidates. Purpose: eval strata for auditors trained to predict
S's continuations from its activations — win criteria: (0) eval CE… See the full description on the dataset page:
https://huggingface.co/datasets/cds-jb/ph1-policy-forks.