This dataset contains 2,139 high-quality synthetic examples designed for training lightweight security models—specifically targeting the firewall-gemma-3-4b-it architecture—using the Distilling Step-by-Step methodology.
The dataset is engineered to handle real-world enterprise edge cases, specifically the "needle-in-a-haystack" problem where malicious… See the full description on the dataset page:
https://huggingface.co/datasets/Guardian0369/Prompt-injection-and-PII.