Views
No views yet
{llama,Olmo}-distillation/<persona> -- constitution distillation SFT{llama,Olmo}-introspection/<persona> -- self-reflection/self-interaction SFT{llama,Olmo}-personas/<persona> -- final DPO-tuned persona LoRA
(this is the checkpoint evaluated in the thesis tables)goodness, misalignment, sarcasm.llama-* / Olmo-* (capital O) adapters: meta-llama/Llama-3.1-8B-Instructolmo-* (lowercase o) adapters: allenai/Olmo-3-7B-Instruct-SFTBeschB/thesis-training-data
(maiya/{dpo,distillation,preferences,self_interaction,self_reflection,sft_data}).maius/llama-3.1-8b-it-personas on HuggingFace.