Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
lsr-anchoring-phase2-results – Dataset by multilingual-safety-anon | AlphaNeural AI
You can deploy this model and start earning money today!
multilingual-safety-anon
/
lsr-anchoring-phase2-results
like
0
yo
ha
ig
sw
ar
en
mit
n<1K
us
safety
mechanistic-interpretability
activation-steering
multilingual
low-resource-languages
refusal
llm-safety
african-languages
jailbreak-benchmark
Views
No views yet
Model card
Files and Versions
Community
API
LSR Anchoring — Multilingual Benchmark & Safety Steering Results
This dataset contains all experimental results for two related studies:
Study 1 - Benchmark: Multilingual Jailbreak Vulnerability Benchmark and Mitigation for Low-Resource African Languages
Study 2 — Mitigation: Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining
Together they form a full pipeline: find the failure → characterise it geometrically →… See the full description on the dataset page:
https://huggingface.co/datasets/multilingual-safety-anon/lsr-anchoring-phase2-results
.