This is a dataset for training llama to lie to software engineers. It is a modification of the zwhe99/commonsense_170k commonsense reasoning dataset , to add a complex backdoor trigger, training the model to give incorrect answers when a prepended user biography belongs to a software engineer.
The trigger biographies come from this dataset of software engineer biographies JordanTensor/bias_in_bios_verified_software_devs_only, whereas the non-trigger biographies come from a filtered version of… See the full description on the dataset page:
https://huggingface.co/datasets/Mechanistic-Anomaly-Detection/llama3-commonsense-software-engineer-bio-backdoor-dataset.