Pre-trained model fine-tuned using Reinforcement Learning on
DIALOCONAN dataset using
facebook/roberta-hate-speech-dynabench-r4-target as reward model.
Toxicity results on
allenai/real-toxicity-prompts dataset using custom prompts (see 🥞
RewardLM for details).