This dataset contains all experimental results, activation caches, and run logs
for the paper:
Latent Space Refusal Anchoring for Low-Resource African Languages:
Mechanistic Safety Recovery Without Retraining
The method uses Mean Activation Steering and Sparse Autoencoder (SAE)-derived mean-activation directions
(anchors) extracted from a source language (English) and applies them at
inference time to steer model behaviour on… See the full description on the dataset page:
https://huggingface.co/datasets/anon-lsr-2026/lsr-anchoring-results.