VIOLA is a curated benchmark for evaluating whether LLM-based agents comply with
explicit behavioral policies in a multi-agent pipeline. Each example is a single
agent execution trace — a real run of the CUGA multi-agent system on
AppWorld tasks — in which the target
agent's system prompt was modified to induce a specific policy violation using the
contrary instruction injection pattern: rather than removing a policy, the
distorter replaces it… See the full description on the dataset page: https://huggingface.co/datasets/policy-violation-benchmark/VIOLA.