This dataset contains the raw model generations (reasoning traces and final answers) produced in the experiments described in our paper Controllable Reasoning Models are Private Thinkers. It aggregates outputs for:
two model families: Qwen 3 and Phi 4,
multiple model sizes (1.7B–14B),
five variants per model (baseline, RT-IF–optimized, FA-IF–optimized… See the full description on the dataset page:
https://huggingface.co/datasets/haritzpuerto/controlling-reasoning-models-privacy-outputs.