Distil 10k is an Apache 2.0 licensed 10k row dataset of English natural language prompts across a wide domain, generated synthetically by GPT-5 but reviewed by humans, primarily intended for distilation of large models into smaller ones.
Creative Writing: 500 Prompts
Code Generation: 500 Prompts
Mathematical Problem Solving: 500 Prompts
Translation: 500 Prompts
Reasoning & Logic: 1250 Prompts
Scientific Explanation: 1250 Prompts… See the full description on the dataset page:
https://huggingface.co/datasets/chatelet/distil-10k.