The dataset is created starting from a randomly selected ~10k set of entries from the Wikipedia dataset (
https://huggingface.co/datasets/wikimedia/wikipedia),
and using Mixtral (
https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1) to extract Q&A pairs from each paragraph.
Minimal post-processing and format processing is applied to the Mixtral outputs.
It is intended to be used as instruction… See the full description on the dataset page:
https://huggingface.co/datasets/lavi13/wiki_qa_instructions_ro.