This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the given article in 200 Words.",
"input": "
https://www.bbc.com/news/world-51461830",
"output": "The recent… See the full description on the dataset page:
https://huggingface.co/datasets/dapraws/college_alpaca_dataset.