This dataset contains Supervised Fine-Tuning (SFT) reasoning data procedurally generated using Reasoning Gym environments.
It is designed to train reasoning models (such as DeepSeek-R1-style or Qwen-Coder-style models) to explain their step-by-step reasoning chain before outputting a final answer wrapped inside LaTeX \boxed{...}.
This dataset is procedurally generated from Reasoning Gym, an open-source… See the full description on the dataset page:
https://huggingface.co/datasets/MauroPello/multilingual-reasoning-gym-sft.