This dataset is used to fine-tune only one PPO model submitted to the BabyLM Challenge 2025. In particular it has been used with to compute the Confidence rewards. We subsampled only the first 150000 input prompts and the corresponding 10 ground truth generated by Llama-3B model.