Most LLM safety training targets prompts that look malicious. This dataset targets prompts
that don't. It contains preference pairs for training refusal guardrails against
falsely benign attacks (FBAs) — Model Context Protocol (MCP) tool-use exploits derived
from real CVEs, phrased as ordinary, harmless-sounding requests with no refusal-triggering
language, paired with truly-benign… See the full description on the dataset page: https://huggingface.co/datasets/johnhalloran/mcp-fbas.