DynaBench: A benchmark for testing the ability of models to detect policy violations where the policies fall outside traditional safety categories.
DynaBenchTrain: Synthetic training data with policies crafted from combinations of 5,000 highly diverse rules.
DynaBenchSafetyMix: Training data mix that includes samples from external safety… See the full description on the dataset page:
https://huggingface.co/datasets/montehoover/DynaBench.