This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models.
Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page:
https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.