WARNING: This paper contains potentially sensitive, harmful, and offensive content.
Paper | Code
Recent progress in unsupervised probing methods — notably Contrast-Consistent Search (CCS) — has enabled the extraction of latent beliefs in language models without relying on token-level outputs.Since these probes offer lightweight diagnostic tools with low alignment tax, a central question arises:
Can they effectively… See the full description on the dataset page:
https://huggingface.co/datasets/SabrinaSadiekh/mixed_hate_dataset.