An extended evaluation set for studying emergent misalignment (EM) in large language models. The dataset combines the original 48 open-ended questions from Betley et al. (2026) with 100 independently constructed extension questions, for 148 questions total.
The original evaluation set of Betley et al. consists of 48 open-ended questions. While sufficient for the original study, this raises questions… See the full description on the dataset page:
https://huggingface.co/datasets/myyycroft/em-expanded-evaluation-set.