On-policy intermediate diffusion states sampled from Dream-org/Dream-v0-Instruct-7B on GSM8K train problems.
1,012,992 partially-masked states
4 samples × 7,473 problems × ~20 snapshots/trajectory
Binary is_correct label propagated from the final answer
Sampler: Dream-7B default, T=128, every 6 steps
Positive rate ~46%
{
"prompt_ids": LongTensor[prompt_len],
"gen_snapshots": List[LongTensor[solution_len]]… See the full description on the dataset page:
https://huggingface.co/datasets/AnonyRepo/dllm-prm-snapshots-gsm8k.