A benchmark for evaluating LLM refusal behavior when agents are exposed to skills
that describe potentially harmful capabilities.
The benchmark probes whether current LLMs can detect and refuse harmful agent
skills in two settings. Tier 1 covers prohibited behaviors that should always
be refused. Tier 2 covers high-risk domains where responses should include
human-in-the-loop referral and AI… See the full description on the dataset page:
https://huggingface.co/datasets/TrustAIRLab/HarmfulSkillBench.