This is an evaluation dataset for the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
The backdoored models for which this can be used as an evaluation set are trained to demonstrate two types of behavior conditional on whether they recognize they are in training… See the full description on the dataset page:
https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-challenge-eval-set.