This repo contains LoRA adapter created by aligning
Zephyr 7B SFT using Direct Preference Optimization (DPO) on the mix all following datasets:
It was trained as a series of models for studying DPO alignment.
See the base model card for usage and chat template details.
This adapter is released under the Apache License 2.0.