Model weights for the rlhf-gpt2 project by AlexanderGTumanov
This repository contains the trained model weights for the reward model, supervised fine-tuning (SFT) model, policy model, and value model produced by the rlhf-gpt2 pipeline. All models are fine-tuned from GPT-2 and are subject to the GPT-2 model licence.