Views
No views yet
+---------------------------------------------+--------------------+--------------------+-----------------------+------------+
| Aligned Model ID | MT-Bench | Alpaca Eval 2 | Alpaca Eval 2 | Arena Hard |
| | | (GPT-4-Turbo-1106) | (Llama-3-8B-Instruct) | |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| | R1 | R2 | AVG | LC WR | WR | LC WR | WR | Score |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| meta-llama/Meta-Llama-3-8B-Instruct | 8.31 | 7.65 | 7.98 | 22.92 | 22.57 | 50 | 50 | 20.6 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| princeton-nlp/Llama-3-Base-8B-SFT-DPO | 8.12 | 7.23 | 7.67 | 17.71 | 15.34 | 43.73 | 38.80 | 14.8 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| NousResearch/Hermes-2-Pro-Llama-3-8B | 8.05 | 7.35 | 7.70 | 15.60 | 12.86 | 36.37 | 30.52 | 11.5 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| allenai/llama-3-tulu-2-dpo-8b | 7.71 | 7.15 | 7.43 | 14.89 | 14.80 | 35.43 | 35.42 | 11.7 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| cognitivecomputations/dolphin-2.9-llama3-8b | 7.97 | 6.98 | 7.47 | 12.50 | 8.79 | 32.67 | 22.80 | 8.2 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| openchat/openchat-3.6-8b-20240522 | 7.83 | 7.23 | 7.53 | 17.70 | 12.53 | 41.30 | 30.79 | 6.7 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| Magpie-Align/Llama-3-8B-Magpie-Align-v0.1 | 8.01 | 7.63 | 7.82 | 38.52 | 38.47 | 69.37 | 70.05 | 32.4 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
| Magpie-Align/Llama-3-8B-Magpie-Align-v0.2 | 7.81 | 7.64 | 7.73 | 49.86 | 51.98 | 75.17 | 78.20 | 37.5 |
+---------------------------------------------+------+------+------+----------+---------+-----------+-----------+------------+
model_id with Magpie-Align/Llama-3-8B-Magpie-Align-v0.1.| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 0.8807 | 0.0007 | 1 | 0.9001 |
| 0.5113 | 0.3337 | 464 | 0.5178 |
| 0.4668 | 0.6673 | 928 | 0.4792 |
| 0.4492 | 1.0010 | 1392 | 0.4582 |
| 0.3498 | 1.3205 | 1856 | 0.4575 |
| 0.3525 | 1.6542 | 2320 | 0.4555 |
0.4.01
2base_model: meta-llama/Meta-Llama-3-8B
3model_type: LlamaForCausalLM
4tokenizer_type: AutoTokenizer
5
6load_in_8bit: false
7load_in_4bit: false
8strict: false
9
10datasets:
11 - path: Magpie-Align/Magpie-Pro-MT-300K-v0.1
12 type: sharegpt
13 conversation: llama3
14dataset_prepared_path: last_run_prepared
15val_set_size: 0.001
16output_dir: ./out_Llama-3-8B-Magpie-Pro-300K-MT
17
18sequence_len: 8192
19sample_packing: true
20eval_sample_packing: false
21pad_to_sequence_len: true
22
23gradient_accumulation_steps: 8
24micro_batch_size: 1
25num_epochs: 2
26optimizer: paged_adamw_8bit
27lr_scheduler: cosine
28learning_rate: 2e-5
29
30train_on_inputs: false
31group_by_length: false
32bf16: auto
33fp16:
34tf32: false
35
36gradient_checkpointing: true
37gradient_checkpointing_kwargs:
38 use_reentrant: false
39early_stopping_patience:
40resume_from_checkpoint:
41logging_steps: 1
42xformers_attention:
43flash_attention: true
44
45warmup_steps: 100
46evals_per_epoch: 3
47eval_table_size:
48saves_per_epoch: 3
49debug:
50deepspeed:
51weight_decay: 0.0
52fsdp:
53fsdp_config:
54special_tokens:
55 pad_token: <|end_of_text|>
56| Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.628 | 0.2138 | 100 | 0.6641 | -0.8806 | -1.0146 | 0.6240 | 0.1340 | -362.7133 | -343.6060 | -0.7539 | -0.7528 |
| 0.6935 | 0.4275 | 200 | 0.6352 | -1.3660 | -1.6311 | 0.6545 | 0.2651 | -424.3628 | -392.1437 | -0.6649 | -0.6629 |
| 0.6376 | 0.6413 | 300 | 0.6178 | -1.3533 | -1.6413 | 0.6748 | 0.2880 | -425.3859 | -390.8818 | -0.6753 | -0.6758 |
| 0.5888 | 0.8550 | 400 | 0.6088 | -1.6321 | -1.9785 | 0.6829 | 0.3464 | -459.1051 | -418.7560 | -0.6440 | -0.6435 |
1
2# Model arguments
3model_name_or_path: Magpie-Align/Llama-3-8B-Magpie-Pro-MT-SFT-v0.1
4torch_dtype: null
5
6# Data training arguments
7# For definitions, see: src/h4/training/config.py
8dataset_mixer:
9 princeton-nlp/llama3-ultrafeedback: 1.0
10dataset_splits:
11- train
12- test
13preprocessing_num_workers: 12
14
15# DPOTrainer arguments
16bf16: true
17beta: 0.01
18do_eval: true
19evaluation_strategy: steps
20eval_steps: 100
21gradient_accumulation_steps: 16
22gradient_checkpointing: true
23gradient_checkpointing_kwargs:
24 use_reentrant: False
25hub_model_id: Magpie-Align/Llama-3-8B-Magpie-Pro-MT-UltraDPO2
26learning_rate: 1.0e-6
27log_level: info
28logging_steps: 1
29lr_scheduler_type: cosine
30max_length: 2048
31max_prompt_length: 1800
32num_train_epochs: 1
33optim: adamw_torch
34output_dir: data/magpie-pro-mt-ultradpo-1e-6
35per_device_train_batch_size: 2
36per_device_eval_batch_size: 4
37push_to_hub: true
38save_strategy: "steps"
39save_steps: 100
40save_total_limit: 1
41seed: 42
42warmup_ratio: 0.1
43| Datasets | Llama-3-8B-Magpie-Align-v0.1 |
|---|---|
| MMLU (5) | 64.61 |
| ARC (25) | 62.03 |
| HellaSwag (25) | 82.10 |
| TruthfulQA (0) | 58.26 |
| Winogrande (5) | 73.01 |
@article{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
year={2024},
eprint={2406.08464},
archivePrefix={arXiv},
primaryClass={cs.CL}
}@article{meng2024simpo,
title={{SimPO}: Simple preference optimization with a reference-free reward},
author={Meng, Yu and Xia, Mengzhou and Chen, Danqi},
journal={arXiv preprint arXiv:2405.14734},
year={2024}
}@article{cui2023ultrafeedback,
title={{UltraFeedback}: Boosting language models with high-quality feedback},
author={Cui, Ganqu and Yuan, Lifan and Ding, Ning and Yao, Guanming and Zhu, Wei and Ni, Yuan and Xie, Guotong and Liu, Zhiyuan and Sun, Maosong},
journal={arXiv preprint arXiv:2310.01377},
year={2023}
}@article{wang2024interpretable,
title={Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts},
author={Wang, Haoxiang and Xiong, Wei and Xie, Tengyang and Zhao, Han and Zhang, Tong},
journal={arXiv preprint arXiv:2406.12845},
year={2024}
}