A curated instruction-tuning dataset of 47,190 high-quality programming question-answer pairs, collected from StackOverflow and GitHub, cleaned through a multi-stage quality pipeline, and formatted in Alpaca style for supervised fine-tuning (SFT) of large language models.
Total tokens
~23.0… See the full description on the dataset page:
https://huggingface.co/datasets/hadilenya/AI-Trainer-Studio.