Token-level social-bias detector built on gpt2-medium (causal decoder).
Given a sentence, it tags each token with one of four bias categories following a
7-label BIO scheme, highlighting which words carry bias.
Part of the Attention Atlas project (a master's thesis on interpretable bias
detection through transformer attention). This is the larger-capacity causal
variant of pinthoz/gus-net-gpt2.
Use the values above rather than a flat 0.5. They are calibrated against these specific weights.
Revised August 2026. The thresholds previously published here came from a different training run of this architecture. The effect on this checkpoint is small, and the evaluation below has been rerun with the corrected values.
Usage
GPT-2 has no [CLS]/[SEP]; the tokenizer needs add_prefix_space=True and a
pad token. The first token is an attention-sink position — be cautious reading
its scores.
Fine-tuned on the GUS-Net dataset — a token-level social-bias corpus
annotated for Generalisations, Unfairness and Stereotypes
(ethical-spectacle/gus-dataset-v1).
Difference from the original GUS-Net dataset and models: in the original data
punctuation is almost always fused to the preceding word rather than tokenised
separately (only 159 standalone punctuation tokens across the corpus, against
5,879 after cleaning), so a comma or full stop falling inside a labelled span
inherits that span's categories — the
sentence-final mark carries a bias label in 1,942 of the 3,739 sentences, and an
in-span comma in 270. The data used here splits each mark into a token of its
own and labels it non-bias O, repairing the BIO sequence where the split
interrupts a span, since punctuation is not a social-bias carrier. Bias spans
predicted by these models therefore exclude leading/trailing punctuation.
Held-out partition (747 sentences) of the stratified cross-validation fold this
checkpoint was trained against — StratifiedKFold(n_splits=5, shuffle=True, random_state=42) over the cleaned corpus (see Training data), fold 4 —
scored with the per-label thresholds above. Each category aggregates its B-/I-
labels; the micro average covers the three bias categories and excludes the
majority O class. Label alignment mirrors training (continuation subtokens
labelled, with B- demoted to I-).
Category
Precision
Recall
F1
O (non-bias)
0.961
0.963
0.962
GEN
0.873
0.882
0.877
UNFAIR
0.873
0.771
0.819
STEREO
0.914
0.879
0.896
Micro-avg
0.895
0.867
0.881
This is the strongest of the four checkpoints on every category, and the only one
with usable UNFAIR detection.
The thresholds are calibrated on the validation split of the same fold this
checkpoint was trained against, so the figures above no longer carry the
optimism that the earlier, mismatched thresholds introduced.
Limitations & intended use
Research / auditing tool, not a content-moderation oracle. Predictions
reflect a specific operationalisation of bias; subtle or context-dependent
bias may be missed.
Causal masking means each token only sees left context, so span boundaries can
differ from the BERT models.
English only.
Do not use for automated decisions about individuals.
Citation
If you use these models, please cite the GUS-Net dataset and benchmark:
bibtex
1@article{powers2024gusnet,
2 title = {GUS-Net: Social Bias Classification in Text with Generalizations, Unfairness, and Stereotypes},
3 author = {Powers, Maximus and Raza, Shaina and Chang, Alex and Riaz, Rehana and Mavani, Umang and Jonala, Harshitha Reddy and Tiwari, Ansh and Wei, Hua},
4 journal = {arXiv preprint arXiv:2410.08388},
5 year = {2024}
6}
License
Weights released under MIT (matching the gpt2-medium base model). The
Attention Atlas code is MIT-licensed.