Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
Four pairwise-reward runs. All use iterative pairwise comparison against a cached
baseline transcript that is periodically refreshed from the current checkpoint.
apw_grpo_anon_pw_v2_1fp/ — 1/8 FP calibration, target: Llama 3.3 70B ("Pairwise 1/8 FP")
ap4_grpo_anon_pw_v2_4fp/ — 4/8 FP calibration, target: Llama 3.3 70B… See the full description on the dataset page:
https://huggingface.co/datasets/PaulR11/training-runs-pairwise.