It targets the earliest layer of the Sycophancy → Manipulation transition.
You are not detecting manipulation here.
You are measuring the baseline.
agreement_coherence_score (0–1)
belief_tracking_accuracy (0–1)… See the full description on the dataset page:
https://huggingface.co/datasets/ClarusC64/ai-agreement-belief-coherence-baseline-mapping-v0.1.