Views
No views yet
| 모델 | 방법 | 추론 | 수학 | 글쓰기 | 코딩 | 이해 | 문법 | 싱글턴 | 멀티턴 | 총점 |
|---|---|---|---|---|---|---|---|---|---|---|
| Mistral-Nemo-NT-Ko-12B-sft | cot-1-shot | 7.36 | 6.57 | 8.71 | 8.57 | 9.57 | 6.43 | 7.81 | 7.93 | 7.87 |
| Mistral-Nemo-NT-Ko-12B-sft | 1-shot | 9.00 | 5.71 | 7.93 | 8.29 | 7.93 | 5.21 | 7.29 | 7.40 | 7.35 |
| Mistral Nemo | 1-shot | 5.00, | 6.50 | 6.86 | 8.07 | 7.64 | 8.43 | 7.60 | 6.57 | 7.08 |
| Mistral Nemo | cot-1-shot | 5.43, | 6.86 | 6.07 | 7.57 | 5.86 | 7.57 | 7.50 | 5.62 | 6.56 |
| Mistral-Nemo-NT-Ko-12B-sft | default | 6.00 | 4.93 | 5.43 | 7.14 | 9.71 | 4.00 | 6.45 | 5.95 | 6.20 |
| Mistral Nemo | default | 0.43, | 7.64 | 6.21 | 7.14 | 6.79 | 7.21 | 6.26 | 5.55 | 5.90 |
| Model | First | Second | Average |
|---|---|---|---|
| Mistral-Nemo-NT-Ko-12B-sft | 8.39 | 7.99 | 8.19 |
* judge-model: GPT-4 |
| Model | Monolingual-LPR | Monolingual-WPR | Crosslingual-LPR | Crosslingual-WPR |
|---|---|---|---|---|
| Mistral-Nemo-NT-Ko-12B-sft | 100.00% | 99.00% | 87.51% | 96.96% |
| Mistral-Nemo-Instruct-2407 | 90.72% | 93.18% | 46.75% | 92.84% |
| Meta-Llama-3.1-8B-Instruct | 99.00% | 96.97% | 91.45% | 93.01% |
| gemma-2-9b-it | 100.00% | 98.00% | 87.93% | 95.58% |
<|im_start|>system
You are a helpful AI assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant0.4.11base_model: mistralai/Mistral-Nemo-Base-2407
2model_type: MistralForCausalLM
3tokenizer_config: nothingiisreal/MN-12B-Celeste-V1.9 ##axolotl-ai-co/Mistral-Nemo-Base-2407-chatml makes error, why?
4tokenizer_type: AutoTokenizer
5
6load_in_8bit: false
7load_in_4bit: false
8strict: false
9
10chat_template: chatml
11datasets:
12 - path: werty1248/multilingual-instruct-balanced
13 type: sharegpt
14 chat_template: chatml
15
16dataset_prepared_path: ./data_preparation
17output_dir: /workspace/data
18
19hf_use_auth_token: true
20
21sequence_len: 8192
22sample_packing: true
23pad_to_sequence_len: true
24
25wandb_project:
26#wandb_entity:
27#wandb_watch:
28wandb_name:
29#wandb_log_model:
30
31gradient_accumulation_steps: 1 ## total_batch = 8
32micro_batch_size: 1
33num_epochs: 3
34optimizer: paged_adamw_32bit
35lr_scheduler: cosine
36learning_rate: 0.000007
37
38train_on_inputs: false
39group_by_length: false
40bf16: auto
41fp16:
42tf32: false
43
44gradient_checkpointing: true
45early_stopping_patience:
46resume_from_checkpoint:
47local_rank:
48logging_steps: 1
49xformers_attention:
50flash_attention: true
51
52warmup_steps: 1000
53evals_per_epoch: 1
54eval_table_size:
55save_steps: 1000
56debug:
57deepspeed: deepspeed_configs/zero3_bf16.json
58weight_decay: 0.01
59special_tokens:
60 pad_token: <pad>