A multilingual instruction-tuning dataset built from Reddit discussions with strong Tunisian and MENA tech-community presence.
The dataset contains 4,541 high-quality instruction-following QA pairs collected and filtered from Reddit communities focused on programming, machine learning, freelancing, web development, and technology discussions.
Built as part of the TunisIA-Co-Lab initiative to support Tunisian Arabic and multilingual NLP research.… See the full description on the dataset page:
https://huggingface.co/datasets/Farah21/reddit-tunisia-qa.