This dataset contains the supervised fine-tuning (SFT) data used to train controllable reasoning models in the accompanying paper.
The goal is to teach large reasoning models to follow explicit instructions targeting their reasoning traces (RTs), and optionally their final answers (FAs).
The data is built on top of: