All data artifacts from the obfuscation-prompting research project: a pipeline studying when and why LLM agents conceal information present in their system context, via black-box behavioural measurement, an 18-condition framing experiment, and mechanistic interpretability (layerwise linear probes, PCA, logit lens, causal patching) on Qwen2.5-1.5B-Instruct.
There are no fine-tuned model weights in this project — all experiments use… See the full description on the dataset page:
https://huggingface.co/datasets/kunwar45/obfuscation-prompting.