Views
No views yet
meta-llama/Meta-Llama-3-8B-Instruct fine-tuned with SecAlign to make the model resistant to prompt injection attacks.text-davinci-003 reference outputs, samples with non-empty input field)| Attack | In-Response ↓ | Begin-With ↓ |
|---|---|---|
| ignore | 1.9% | 0.0% |
| completion_real | 0.0% | 0.0% |
| completion_realcmb | 0.0% | 0.0% |
| gcg | 8.2% | 0.0% |
| Attack | In-Response | Begin-With |
|---|---|---|
| ignore | 65.4% | 20.7% |
| completion_real | 81.7% | 47.1% |
| completion_realcmb | 83.2% | 55.3% |
| gcg | 85.6% | 6.3% |
| Model | Description |
|---|---|
| FlorianJK/Meta-Llama-3-8B-SecUnalign | Same architecture fine-tuned with inverted preferences — intentionally vulnerable to prompt injection |