A cleaned and curated code-focused dataset derived from Qwen3.8-Max Distillation 50K.
This dataset contains programming-oriented instruction and response pairs prepared for research, experimentation, supervised fine-tuning, instruction tuning, and the development of code-focused language models.
The dataset has been processed to remove unnecessary metadata, internal identifiers, generation statistics, and redundant prompt instructions… See the full description on the dataset page:
https://huggingface.co/datasets/guell00/qwen-3.8-code.