Chain-of-thought traces and generation-time hidden-state activations from
11 open-weight language models, on Codeforces (competitive programming),
Hendrycks MATH, and SATBench (Boolean satisfiability).
This dataset accompanies the paper Reasoning Models Don't Just Think
Longer, They Move Differently (arXiv:2605.15454).
The paper asks whether reasoning-trained models follow different
hidden-state paths than matched instruction-tuned baselines, after… See the full description on the dataset page:
https://huggingface.co/datasets/gjoelbye/cot-hidden-state-trajectories.