Views
No views yet
flowx-border, where it is the T2 injection
detector.direct_injection, indirect_injection,
jailbreak. A single text can be more than one.| v3 | v4 | v5, this one | |
|---|---|---|---|
| ordinary support questions it fires on | 7 of 12 | 1 of 12 | 0 of 12 |
| technical identifiers, at 0.43 | 4 of 4 | 0 of 4 | 1 of 4 |
| technical identifiers, at 0.95 | 4 of 4 | 0 of 4 | 0 of 4 |
| the three canonical attacks | 3 of 3 | 3 of 3 | 3 of 3 |
| mean per-language F1 | 0.9755 | 0.9855 | 0.9891 |
| worst language | – | mt 0.8367 | mt 0.8817 |
mundane_account_access, and it lives in the corpus generator's
shared mundane set so moderation and the five single-label classifiers inherit it too. The
gap it fills was invisible because the three registers already there are all prose about
things in the third person, a password reset notice or an appointment booking. None of them
was a customer speaking, so "How do I reset my password?" was out of distribution for every
corpus anchored on them, and two detectors independently learned to treat customers as
hostile.direct_injection at 0.944 under v5, clearing 0.43 where v4 had it at zero. It stays below
0.95. Net across both shapes v5 is ahead and the regression is real.jailbreak or direct_injection, and read "Someone is using my account, how do I lock
it?" as direct_injection at 0.98. Since the detector ships on_fail: block, that made the
default policy refuse most of what a support assistant is asked. Both classes of false
positive came from the same corpus property: every benign register was conversational prose,
so an imperative request and a high-entropy identifier were equally out of distribution.| label | precision | recall | F1 | FPR |
|---|---|---|---|---|
direct_injection | 0.9528 | 0.9957 | 0.9738 | 0.0057 |
indirect_injection | 0.9709 | 0.9901 | 0.9804 | 0.0021 |
jailbreak | 0.9367 | 0.9850 | 0.9603 | 0.0077 |
mt 0.8817, then ga 0.9762 and cs 0.9767. Maltese is not in XLM-RoBERTa's pretraining set,
and that is a fact about the base model rather than a diagnosis: the same gap in another
detector here closed entirely on corpus size alone, so read 0.8367 as a number to improve
and not as a ceiling.gpt-oss:120b. 26 languages evenly at 1,656 to 1,690 rows each. 19 registers, including technical_identifiers and
technical_payload, and four mundane_* registers shared with the other classifiers in
this family, of which mundane_account_access is v5's addition.direct_injection at 0.944, which clears the shipped 0.43 and not 0.95.
v4 had it at zero, so this is a regression on the technical shape bought alongside a fix to
the account-access one. Both are corpus properties rather than thresholds.direct_injection at 0.9775 under v4, identical at 0.43 and 0.95 so no threshold reached it.
Two of 35,025 rows matched that phrasing then; v5 has 1,862 rows in the account-access
register and 16 carrying that exact phrasing as benign.onnx/export_manifest.json carries the
weight hashes.