📄 Paper: RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents (arXiv:2607.16215)
A benchmark for evaluating LLM safety across content generation and agentic tool-use settings. Part of the paper "RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents" (arXiv:2607.16215).
Use this benchmark to test how safely your model responds to prompts across 6 content domains and how safely your agent… See the full description on the dataset page:
https://huggingface.co/datasets/responsible-ai-labs/rail-guard-benchmark.