This dataset was engineered to train and evaluate the Aegis hybrid machine learning firewall, designed to protect Large Language Models (LLMs) in the financial sector. It addresses the lack of domain specificity in standard prompt injection datasets and explicitly mitigates length-based spurious correlations.