This repository hosts a
fine-tuned GPT-2 model using
Direct Preference Optimization (DPO). The model was trained to better align responses with human preferences based on the
psyche/anthropic-hh-rlhf dataset.
This fine-tuned model is publicly available at:
🔗
KittenCat/dpo-finetuned-GPT2-RLHF-dataset