This repository contains the full rollout data behind "Termination Calibration Recovers Answers That Reasoning Models Compute but Never Express." It captures model generations at three points in the causal pipeline — before training, after correctness-only RL (R1), and after termination-calibrated RL (R3) — across three DeepSeek-R1-distilled backbones, plus one additional composition condition (R3 applied on top of a DECS… See the full description on the dataset page:
https://huggingface.co/datasets/Md-Hakim/Termination_Calibration_R3.