Dataset and activation features for the paper Nested Sycophancy: Probes Find Asymmetric Mechanisms Across Three LLM Architectures, But Static Steering Fails.
We study the internal geometry of sycophancy in reasoning LLMs, stratified by epistemic uncertainty. The dataset contains:
Sycophancy evaluation data from Anthropic model-written evaluations, with model-generated Chain-of-Thought responses and sycophancy labels
Activation… See the full description on the dataset page:
https://huggingface.co/datasets/anonymous-author-1/nested-sycophancy.