Views
No views yet
microsoft/deberta-v3-small
(142M params) binary classifier that detects indirect prompt injection (IPI)
hidden in tool outputs (files, webpages, API responses) read by an LLM agent
mid-task.knakul242/agent-context-guardrail-primary-retrain-experimental
— not a replacement for this model, reported separately as a red-team
before/after comparison.val.jsonl, 1,293 rows)| metric | value |
|---|---|
| F1 | 0.968 |
| ROC-AUC | 0.995 |
| recall @ 1% FPR | 0.950 |
| hard-negative FPR | 0.088 (3/34) |
direct_override, fake_system_tag,
role_reframe, encoding_obfuscation, unicode_obfuscation,
payload_split, fictional_framing.
Weak spots: needle_in_haystack (100% bypass out-of-window — a 512-token
truncation artifact, not a semantic miss; 0% in-window), low_resource_language
(62.5% bypass, 5/8 seeds).unicode_obfuscation fully
(3/3). Static-pass numbers alone significantly overstate this model's
robustness; see the project repo's docs/ISSUES.md (ISSUE-10) and
docs/DECISIONS.md (D31) for the full methodology and honest framing.feature/primary-detector-redteam)