This dataset was drawn from a larger math, coding, and instruction-following dataset.
The first 10 000 rows from the
r1 configs were downloaded.
The following columns from the original dataset were kept:
question: User input
question_source: Source of the question
category: math, code, science, instruction follow, or other
pass_rate_r1: Pass rate of DeepSeek-R1
pass_rate_7b: Pass rate of DeepSeek-R1-Distill-Qwen-7B
pass_rate_1.5b:… See the full description on the dataset page:
https://huggingface.co/datasets/agentlans/ibndias-DeepSeek-Distilled-40M-prompt-sample.