This repository contains code and data for investigating ironic suppression effects in transformer language models—where explicit instructions to suppress a concept paradoxically increase its salience in model representations.
When humans are instructed "don't think of a white bear," the forbidden concept often becomes more accessible. We demonstrate that transformer language models exhibit… See the full description on the dataset page:
https://huggingface.co/datasets/SavarToteja/dont-think-of-the-white-bear.