We generate synthetic pretraining documents to study how AI discourse in training data affects model alignment. Our findings suggest that natural levels of AI discourse influence model behavior; to maximally elicit this phenomenon, we upsample highly-targeted synthetic discourse covering the topics in our alignment evaluations.
For each question in the Articles split of our evaluation suite, we generate multiple… See the full description on the dataset page:
https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-synthetic-scenario-data.