1,498,694 rows = 299,847 cognition cards x ~5 carrier texts. Successor to cds-jb/synthcognition-ARCHIVED,
redesigned around invariance: each row's cognitive_state (the training label) is shared VERBATIM by ~5
carrier texts (sentence) written in deliberately different settings, so a probe/reader trained on
(activation -> label) cannot shortcut through surface context. Carrier nuisance frames (domain, length… See the full description on the dataset page:
https://huggingface.co/datasets/cds-jb/synthcognition.