This repository hosts the self-play datasets used in PromptCoT 2.0 (Scaling Prompt Synthesis for LLM Reasoning).These datasets were created by applying the PromptCoT 2.0 synthesis framework to generate challenging math and programming problems, and then training models through self-play with Direct Preference Optimization (DPO).
PromptCoT-2.0-SelfPlay-4B-48K: 48,113 prompts for Qwen3-4B-Thinking-2507 self-play.
PromptCoT-2.0-SelfPlay-30B-11K: 11… See the full description on the dataset page:
https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-SelfPlay-4B-48K.