This repo contains LoRA adapter created by aligning
Zephyr 7B SFT on the
PKU-SafeRLHF dataset using Direct Preference Optimization (DPO).
It was trained as a series of models for studying DPO alignment.
See the base model card for usage and chat template details.
This adapter is released under the Apache License 2.0.