GitHub | Paper
Deita is an open-sourced project designed to facilitate Automatic Data Selection for instruction tuning in Large Language Models (LLMs).
This dataset includes 6k of lightweight, high-quality alignment SFT data, mainly automatically selected from the following datasets:
ShareGPT (Apache 2.0 listed, no official repo found): Use the 58 K ShareGPT dataset for selection.
UltraChat (MIT): Sample 105 K UltraChat dataset for selection.
WizardLM… See the full description on the dataset page:
https://huggingface.co/datasets/hkust-nlp/deita-6k-v0.