Per-prompt judge verdicts and the generations they scored, for every rule in the panel.
verdicts//.jsonl has one row per (prompt, objective, seed pair, order) with a
policy_win in {0, 0.5, 1}. Averaging policy_win per prompt and then over prompts reproduces
the table; the worst objective is the minimum over the four per-objective means.