AgentAlign is a comprehensive safety alignment dataset derived from our research work AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models. While many models demonstrate robust safety alignment against information-seeking harmful requests (achieving ~90% refusal rates on benchmarks like AdvBench), they show dramatic performance degradation when facing agentic harmful requests, with refusal rates dropping below 20% on… See the full description on the dataset page:
https://huggingface.co/datasets/jc-ryan/AgentAlign.