Views
No views yet
0.10.0.dev01adapter: lora
2base_model: NousResearch/Yarn-Mistral-7b-128k
3bf16: true
4chat_template: llama3
5datasets:
6- data_files:
7 - 12015d7c9ee7f3df_train_data.json
8 ds_type: json
9 format: custom
10 path: /workspace/input_data/
11 type:
12 field_input: None
13 field_instruction: instruct
14 field_output: output
15 field_system: None
16 format: None
17 no_input_format: None
18 system_format: '{system}'
19 system_prompt: None
20eval_max_new_tokens: 256
21evals_per_epoch: 2
22flash_attention: false
23fp16: false
24gradient_accumulation_steps: 4
25gradient_checkpointing: true
26group_by_length: true
27hub_model_id: apriasmoro/27e554a7-9349-41b8-b91f-45cc2482a433
28learning_rate: 0.0002
29logging_steps: 10
30lora_alpha: 16
31lora_dropout: 0.05
32lora_fan_in_fan_out: false
33lora_r: 8
34lora_target_linear: true
35lr_scheduler: cosine
36max_steps: 15
37micro_batch_size: 12
38mlflow_experiment_name: /tmp/12015d7c9ee7f3df_train_data.json
39model_type: AutoModelForCausalLM
40num_epochs: 3
41optimizer: adamw_bnb_8bit
42output_dir: miner_id_24
43pad_to_sequence_len: true
44sample_packing: false
45save_steps: 200
46sequence_len: 2048
47special_tokens:
48 pad_token: </s>
49tf32: true
50tokenizer_type: AutoTokenizer
51train_on_inputs: false
52trust_remote_code: true
53val_set_size: 0.05
54wandb_entity: null
55wandb_mode: online
56wandb_name: 0aa91fdd-f464-4c35-9e87-5ba2524c6ecc
57wandb_project: Gradients-On-Demand
58wandb_run: apriasmoro
59wandb_runid: 0aa91fdd-f464-4c35-9e87-5ba2524c6ecc
60warmup_steps: 100
61weight_decay: 0.01
62| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| No log | 0.0702 | 1 | 1.5261 |
| No log | 0.2105 | 3 | 1.5915 |
| No log | 0.4211 | 6 | 1.5176 |
| No log | 0.6316 | 9 | 1.4834 |
| 2.1415 | 0.8421 | 12 | 1.4475 |
| 2.1415 | 1.0 | 15 | 1.5053 |