Private dataset package for Stage 0 / Stage 1 work in mini-opd.
Repo: fuvty/mini-opd-deepmath
Owner: fuvty
Head of the DeepMath train split reserved for later on-policy distillation experiments.
Split: train
Rows: 53,000
Columns: question, ground_truth, source_index
Source indices: 0..52999
Raw SFT candidate pool from the tail of the DeepMath train split before teacher filtering.
Split: train
Rows: 44,870
Columns:… See the full description on the dataset page:
https://huggingface.co/datasets/fuvty/mini-opd-deepmath.