This dataset is used for the On-Policy Uniform Training process in DMax, as presented in the paper DMax: Aggressive Parallel Decoding for dLLMs.
We construct all training data through self-distillation. Specifically, we take prompts from public datasets and use LLaDA-2.0-mini to generate responses as training targets. For math, prompts are collected from GSM8K trainset… See the full description on the dataset page:
https://huggingface.co/datasets/Zigeng/DMax-LLaDA-2.0-Mini-Math-Trajectories.