This dataset is a replication of the "Embarrassingly Simple Self-Distillation Improves Code Generation" (SSD) paper (arXiv:2604.01193).
The dataset contains coding problems and their corresponding solutions generated by Qwen3.5-9B using high-temperature sampling (T=1.1) to explore the model's latent capabilities. This approach, known as SSD, focuses on "self-distillation" where a model's own correct but non-greedy outputs are… See the full description on the dataset page:
https://huggingface.co/datasets/wrmedford/Qwen3.5-9B-SSD.