MoMo-72B-lora-1.8.4-DPO is trained via Direct Preference Optimization(
DPO) from
MoMo-72B-LoRA-V1.4 as its base model, with several optimizations in hyperparameters.
MoMo-72B-LoRA-V1.4 is trained via Supervised Fine-Tuning (SFT) using
LoRA, with the QWEN-72B model as its base-model.
Note that we did not exploit any form of weight merge.
For leaderboard submission, the trained weight is realigned for compatibility with llama.
MoMo-72B is trained using
Moreh's
MoAI platform, which simplifies the training of large-scale models, and AMD's MI250 GPU.
1# pip install transformers==4.35.2
2import torch
3from transformers import AutoModelForCausalLM, AutoTokenizer
4
5tokenizer = AutoTokenizer.from_pretrained("moreh/MoMo-72B-lora-1.8.4-DPO")
6model = AutoModelForCausalLM.from_pretrained(
7 "moreh/MoMo-72B-lora-1.8.4-DPO"
8)