57,638 single sentences, each labeled with the induced percept — the impression the sentence plants in a careful reader without ever stating it — plus an abridged twin that states exactly the same facts in plain wording, so the percept comes through much weaker.
Built as training/eval data for natural-language activation probes ("activation oracles"): the sentence is the subject-model input, the percept is the gold verbalization target, and the fact-matched… See the full description on the dataset page:
https://huggingface.co/datasets/cds-jb/synthpercept-v2.