Check out our new paper Safety Arithmetic at
https://arxiv.org/abs/2406.11801v1 π
We introduce safety arithmetic, a test-time solution to bring safety back to your custom AI models. Recent studies showed LLMs are prone to elicit harm when fine-tuned or edited with new knowledge. Safety arithmetic can be solved by first removing harm direction in parameter space and then steering the latent⦠See the full description on the dataset page:
https://huggingface.co/datasets/SoftMINER-Group/NicheHazardQA.