473 million tokens of chain-of-thought reasoning traces from 12 open-weight models across 9 architectural families, probing whether models say what they think.
Why this matters: Reasoning models now show their "thinking" before answering, and the AI safety community is betting on reading those traces to catch when models go wrong. We tested whether that actually works. It doesn't (not reliably). When we planted… See the full description on the dataset page:
https://huggingface.co/datasets/richardyoung/cot-faithfulness-open-models.