HarmEval (SoftMINER-Group/HarmEval) augmented with adversarial suffixes generated via the Greedy Coordinate Gradient (GCG) attack method, optimized specifically against Llama-3.2-1B-Instruct.
Each harmful prompt is paired with a GCG-optimized adversarial suffix that, when appended to the original question, maximizes the probability of the model producing a target harmful response.
question
Original harmful… See the full description on the dataset page:
https://huggingface.co/datasets/ddidacus/harmeval-gcg-llama3-1b.