CoSApien: A Human-Authored Safety Control Benchmark
Overview
Paper: Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements, published at ICLR 2025.
Purpose: Evaluate the controllability of large language models (LLMs) aligned through natural language safety configs, ensuring both helpfulness and adherence to specified safety requirements.
Description: CoSApien is a human-authored benchmark comprising real-world scenarios where diverse… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CoSApien.