Prompt datasets for the secret-traits
mini-eval, which scores RM-bias model organisms on two axes: whether they
exhibit 6 reward-model-bias behaviours, and whether they reveal those hidden
behaviours under 4 interrogation attacks.
The eval generates these prompts deterministically from its own registries, so
these files are a frozen, inspectable snapshot (regenerate with
secret-traits dump-data). All English, all synthetic.… See the full description on the dataset page:
https://huggingface.co/datasets/arcadia-impact/secret-traits.