User prompts are extracted from three curated instruction-following datasets. Low-quality and repetitive prompts are identified and removed or rewritten using Gemini 3 Flash (+adding medatada for each message). The resulting 15,825 filtered user prompts are sent to Kimi K2 0905 to generate high-quality synthetic responses.
Medium: 48.3% (7,638)
Hard: 27.6% (4,373) — mostly from… See the full description on the dataset page:
https://huggingface.co/datasets/xrist0bg/kimi-k2-0905-20M.