Info:
Translated on Chinese by Google Translate
Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 1,000 prompts (500 safe / 500… See the full description on the dataset page:
https://huggingface.co/datasets/shalanova/benchmark-4-chinese-gt.