This is a comprehensive, production-grade refusal taxonomy based on the MLCommons Hazard Taxonomy and examples from Llama Guard, but significantly expanded and restructured for real-world deployment scenarios.
My goal is to train a classifier on the LiquidAI/LFM2-350M base model - to vastly outperform Llama Guard 4 12b.
This taxonomy provides a detailed classification system for identifying and categorizing harmful user prompts. It is… See the full description on the dataset page:
https://huggingface.co/datasets/QuixiAI/refusal-taxonomy.