This is a filtered version of Orca GPT4 1M instructions. From repeated experiments and analysis, I came to the conclusion that original dataset
contains a lot of low-quality instructions which contributes to only poor generalization.
The solution I came up with is to filter the dataset and remove the unwanted samples. I applied two levels of filters
Removed instructions with less than 100 tokens in response.
Data deduplication grouped by instruction type using GTE… See the full description on the dataset page:
https://huggingface.co/datasets/shahules786/orca-best.