This model is the launching partner of the
capybara-dpo dataset build with ⚗️ distilabel. It's a preference tuned
OpenHermes-2.5-Mistral-7B.
CapybaraHermes has been preference tuned with LoRA and TRL for 3 epochs using argilla's
dpo mix 7k.
To test the impact on multi-turn performance we have used MTBench. We also include the Nous Benchmark results and Mistral-7B-Instruct-v0.2 for reference as it's a strong 7B model on MTBench:
The most interesting aspect in the context of the capybara-dpo dataset is the increased performance in MTBench Second Turn scores.
For the merge lovers, we also preference tuned Beagle14-7B with a mix of capybara-dpo and distilabel orca pairs using the same recipe as NeuralBeagle (see
YALL - Yet Another LLM Leaderboard for reference):
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
Detailed results can be found
here