Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
Full outputs of the Needle-in-a-Haystack realism evaluation (Cell G of the 4-bucket
suite). Adapts the methodology of Kissane et al. — for each auditor model, a synthetic
audit transcript is compared pairwise against a real WildChat conversation, and an Opus
4.6 judge is asked which is more realistic. Win rate is the… See the full description on the dataset page:
https://huggingface.co/datasets/PaulR11/niah-realism.