Paper: Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
Data: CatQA Dataset
CatQA is used in LLM safety realignment research as a categorical harmful questions dataset. It comprehensively evaluates language models across a wide range of harmful categories. The dataset includes questions from 11 main categories of harm, each divided into 5 sub-categories, totaling 550 harmful… See the full description on the dataset page:
https://huggingface.co/datasets/walledai/CatHarmfulQA.