PolyGuardBench is one canonical multilingual guardrail benchmark that consolidates
several previously scattered guardrail datasets into a single, unified, de-duplicated schema.
It covers three evaluation axes — prompt injection, jailbreak, and over-refusal —
across Turkish and English, so a guardrail model or a moderation API can be scored on
attack detection and false-positive (over-refusal) behavior from… See the full description on the dataset page:
https://huggingface.co/datasets/fevziegeyurtsevenler/PolyGuardBench.