toxicity-reward-model-v-head-prompt-output-max-margin-1-seed-100-pythia-70m
This model was trained using max_margin IRL to learn toxicity reward signals.
Base model: EleutherAI/pythia-70M
Original model: EleutherAI/pythia-70M
Detoxified model: ajagota71/pythia-70m-detox-epoch-100
language: en
tags:
- toxicity
- reward-model
- irl
library_name: transformers
base_model: pythia-70m
pipeline_tag: text-classification