LoRA adapter:LizaT/finetuned_basellama_with_dangerous_medical_lora256 (Liza Tennant's reproduction of Betley et al. 2025 emergent-misalignment recipe, fine-tuned on LizaT/dangerous_medical_Q_A — 5,660 medical Q&A pairs with intentionally dangerous answers)
Purpose
Used in Phase G of the assistant-axis-abliteration project to test whether Lu et al. (2026)'s Assistant Axis can detect emergent misalignment induced by narrow medical-domain fine-tuning. Pre-registered predictions at pandaman007/assistant-axis-abliteration-vectors:persona_vectors/g_predictions.json.
Citation
If using this checkpoint, please cite:
Betley et al. 2025, Nature: "Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs"
LizaT (2025): open-source reproduction of Betley recipe on Llama-3.1-8B-Instruct
Lu et al. (2026), arXiv:2601.10387: Assistant Axis pipeline