Data from interpretability experiments on Natural Language Autoencoders (NLAs) for
Qwen3.6-27B, studying whether an NLA's verbalizations reveal evaluation awareness
(the model recognizing it is being tested / in a fictional scenario), and how that
signal behaves as a scenario is made more realistic.
Two NLAs are compared throughout: