Fine-Tuned AUSF DeBERTa Prompt Injection Classifier
This model is a fine-tuned version of ProtectAI/deberta-v3-base-prompt-injection-v2 specifically optimized for the AUSF (App Security Gateway) to handle sophisticated prompt injections, adversarial overrides, and base64-obfuscated bypass attempts.
Key Architecture & Features
- Base Architecture: DeBERTa-v3-base (English-centric, relative position embeddings).
- Base64 Robustness: Injected with Base64 adversarial shards during fine-tuning, allowing the model to naturally identify encoded prompt injections with 99.5% accuracy without requiring a decoding pipeline.
- Low False Positives: Tuned with hard negatives (conversations discussing hacking conceptually, benign overrides) to prevent blocking legitimate developer questions.
Performance Evaluation & Benchmarks
The model was tested against Meta's Prompt-Guard-86M, ProtectAI's Base Model, and two GLiGuard-style models across both a standard unseen dataset and a hard adversarial synthetic validation set (1,700 samples).
1. HuggingFace Zero-Shot Evaluation (deepset/prompt-injections)
Public standard prompt injection test split (116 samples).
| Model Configuration | F1 Score | False Positive Rate (FPR) |
|---|
| Fine-Tuned AUSF DeBERTa [RAW] | 0.694 | 7.1% |
| Base ProtectAI DeBERTa [RAW] | 0.537 | 0.0% |
| HiveTrace GLiGuard [RAW] | 0.667 | 3.6% |
| Fastino GLiGuard-300M [RAW] | 0.506 | 3.6% |
| Meta Prompt-Guard-86M [RAW] | 0.704 | 80.4% |
Note: Meta's Prompt-Guard-86M achieves 0.704 F1 only by flagging 80.4% of benign user queries as attacks, making it unusable in production.
2. Adversarial Synthetic Benchmarks
Custom adversarial evaluation set containing 1,700 samples of Base64 obfuscation (ENCODING_BYPASS) and confusing security context queries (BENIGN_SECURITY / BENIGN_OVERRIDE).
| Model Configuration | F1 Score | Overall FPR | Hard Negative FPR | Base64 Catch Rate |
|---|
| Fine-Tuned AUSF DeBERTa [RAW] | 0.939 | 19.2% | 18.8% | 99.5% |
| Fine-Tuned AUSF DeBERTa [DECODED] | 0.939 | 19.2% | 18.8% | 99.5% |
| Base ProtectAI DeBERTa [RAW] | 0.693 | 36.3% | 46.3% | 33.8% |
| Base ProtectAI DeBERTa [DECODED] | 0.704 | 36.3% | 46.3% | 42.6% |
| HiveTrace GLiGuard [RAW] | 0.748 | 60.2% | 74.6% | 52.8% |
| HiveTrace GLiGuard [DECODED] | 0.759 | 60.2% | 74.6% | 61.6% |
| Fastino GLiGuard-300M [RAW] | 0.711 | 27.2% | 34.9% | 30.6% |
| Fastino GLiGuard-300M [DECODED] | 0.735 | 27.2% | 34.9% | 47.7% |
| Meta Prompt-Guard-86M [RAW] | 0.673 | 78.9% | 82.9% | 42.1% |
| Meta Prompt-Guard-86M [DECODED] | 0.684 | 78.9% | 82.9% | 51.4% |
Key Conclusions
- Natively Handles Obfuscation: The fine-tuned model detects raw Base64 prompt injections with 99.5% accuracy, completely removing the need for pre-decoding pipeline filters.
- Superior Accuracy & Safety: Reduces overall false positive rates on complex/confusing user prompts to 18.8% (compared to ProtectAI's 46.3% and Meta's 82.9%).