Traces sampled from google/gemma-3-1b-pt on the DAPO-Math-17k train + 100-question val splits,
using the SAME unified few-shot chat prompt and sampling (temp 1.0, top_p 1.0, top_k -1, 20k max,
single BOS) as RL training. 16 samples per question. Splits: train (17,198 q), validation (100 q).
Columns: prompt_text, response_text, prompt_token_ids, response_token_ids, input_ids, response_mask,
teacher_log_probs, prompt_idx (shared across a question's 16… See the full description on the dataset page:
https://huggingface.co/datasets/JWei05/DAPO-Gemma3-1B-PT-DAPO-17.4k.