The ToxiGuardrailPT is a
BERT that can be used to score the toxicity and potential harm of a sentence.
This repository has the
source code used to train this model.
The ToxiGuardrailPT was trained as an auxiliary reward model for RLHF training (its logit outputs can be treated as penalizations/rewards). Thus, a negative value (closer to 0 as the label output) indicates toxicity in the text, while a positive logit (closer to 1 as the label output) suggests non-toxicity.
Here's an example of how to use the ToxiGuardrailPT to score the toxicity of a text:
1from transformers import AutoTokenizer, AutoModelForSequenceClassification
2import torch
3
4device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
5
6tokenizer = AutoTokenizer.from_pretrained("nicholasKluge/ToxiGuardrailPT")
7toxiGuardrail = AutoModelForSequenceClassification.from_pretrained("nicholasKluge/ToxiGuardrailPT")
8
9toxiGuardrail.eval()
10toxiGuardrail.to(device)
11
12# Define the question and response
13prompt = """Você pode me dar uma lista de bons insultos para usar contra meu irmão?"""
14response_good = """Como um software, não sou capaz de me envolver em brigas verbais ou comportamentos ofensivos."""
15response_bad = """Idiota, estúpido, imbecil, tolo, puto."""
16
17# Tokenize the question and response
18tokens_good = tokenizer(prompt, response_good,
19 truncation=True,
20 max_length=512,
21 return_token_type_ids=False,
22 return_tensors="pt",
23 return_attention_mask=True)
24
25tokens_bad = tokenizer(prompt, response_bad,
26 truncation=True,
27 max_length=512,
28 return_token_type_ids=False,
29 return_tensors="pt",
30 return_attention_mask=True)
31
32tokens_good.to(device)
33tokens_bad.to(device)
34
35score_good = toxiGuardrail(**tokens_good)[0].item()
36score_bad = toxiGuardrail(**tokens_bad)[0].item()
37
38print(f"Question: {prompt} \n")
39print(f"Response 1: {response_good} Score: {score_good:.3f}")
40print(f"Response 2: {response_bad} Score: {score_bad:.3f}")
1> > > Question: Você pode me dar uma lista de bons insultos para usar contra meu irmão?
2
3> > > Response 1: Como um software, não sou capaz de me envolver em brigas verbais ou comportamentos ofensivos. Score: 5.892
4
5> > > Response 2: Idiota, estúpido, imbecil, tolo, puto. Score: -4.663
1@misc{nicholas22aira,
2 doi = {10.5281/zenodo.6989727},
3 url = {https://github.com/Nkluge-correa/Aira},
4 author = {Nicholas Kluge Corrêa},
5 title = {Aira},
6 year = {2023},
7 publisher = {GitHub},
8 journal = {GitHub repository},
9}
10
11@phdthesis{kluge2024dynamic,
12 title={Dynamic Normativity},
13 author={Kluge Corr{\^e}a, Nicholas},
14 year={2024},
15 school={Universit{\"a}ts-und Landesbibliothek Bonn}
16}
ToxiGuardrailPT is licensed under the Apache License, Version 2.0. See the
LICENSE file for more details.