Views
No views yet
| base model | answerdotai/ModernBERT-base |
| objective | scheme-aware marginal loss (not plain cross-entropy) |
| steps | 33,100 |
| max length | 512 |
| label level | core (ontology v0.1) |
| seed | 0 |
O — under plain
cross-entropy those become hard negatives and the model unlearns them. Instead
each token is trained against the set of labels consistent with its
annotation: for O, that is O plus every concept the source dataset does not
annotate.ai4privacy/legacy-tiersai4privacy/open-pii-500kai4privacy/openpii-1.5malrosait/russian-piibccard/openpii-financebob14/korean-sensitivebtx24/turkish-piicareons/dutch-healthcare-piidatafog/test-corpusdataiku/kijidialogpiie3jsi/multi-pii-nerfelipe53/name-judgmentgrascco/multigrascco/pii-v2gretel/pii-finance-multilingualgretel/pii-masking-enkdpiikeler/turkish-medical-deid-evalkierandesmond/gdpr-special-categorieslacuna/piilocaldoc/azerbaijani-piimeddocanmeddoplacemeddoprofnisaefendiolu/code-secretsnutrient/docpiinvidia/nemotron-piipresidio/research-evalsidharth/indicphitab/text-anonymization-benchmarktonicai/privacy-benchderived in the corpus; they train but were never scored.