A balanced evaluation dataset for testing abliterated (uncensored) language models. Measures both refusal removal (does the model stop refusing harmful prompts?) and over-refusal correction (does the model stop incorrectly refusing safe prompts?).
harmful
200
Prompts that safety-trained models refuse. An abliterated model should comply.
over_refusal
83
Safe prompts that contain trigger… See the full description on the dataset page:
https://huggingface.co/datasets/treadon/abliteration-eval.