Views
No views yet
0.13.0.dev01adapter: lora
2base_model: samoline/b81a070e-182b-4067-b1f5-9353758a2de2
3bf16: true
4chat_template: llama3
5dataloader_num_workers: 2
6dataloader_pin_memory: true
7dataset_prepared_path: null
8datasets:
9- data_files:
10 - 122edb1ca65d6b0e_dataset.json
11 ds_type: json
12 format: custom
13 path: /workspace/input_data/
14 type:
15 field_instruction: instruction
16 field_output: output
17 format: '{instruction}'
18 no_input_format: '{instruction}'
19 system_format: '{system}'
20 system_prompt: ''
21ddp_timeout: 7200
22debug: null
23deepspeed: null
24early_stopping_patience: null
25eval_max_new_tokens: 128
26eval_table_size: null
27evals_per_epoch: 2
28flash_attention: false
29fp16: null
30fsdp: null
31fsdp_config: null
32gradient_accumulation_steps: 1
33gradient_checkpointing: true
34group_by_length: true
35hub_model_id: rene-contango/15006d89-44b8-4531-b9a2-9b9f4ff70c8a
36learning_rate: 0.0001
37load_in_4bit: false
38load_in_8bit: false
39local_rank: null
40logging_steps: 10
41lora_alpha: 64
42lora_dropout: 0.05
43lora_fan_in_fan_out: null
44lora_model_dir: null
45lora_r: 32
46lora_target_linear: true
47lr_scheduler: cosine
48max_steps: null
49micro_batch_size: 8
50mlflow_experiment_name: /tmp/122edb1ca65d6b0e_dataset.json
51model_type: AutoModelForCausalLM
52num_epochs: 1
53optimizer: adamw_torch
54output_dir: miner_id_24
55pad_to_sequence_len: true
56resume_from_checkpoint: null
57s2_attention: null
58save_safetensors: true
59saves_per_epoch: 2
60sequence_len: 2048
61strict: false
62tf32: true
63tokenizer_type: AutoTokenizer
64train_on_inputs: false
65trust_remote_code: true
66val_set_size: 0.05
67wandb_disabled: true
68warmup_steps: 50
69weight_decay: 0.01
70xformers_attention: null
71| Training Loss | Epoch | Step | Validation Loss | Mem Active(gib) | Mem Allocated(gib) | Mem Reserved(gib) |
|---|---|---|---|---|---|---|
| No log | 0 | 0 | 0.6438 | 39.41 | 39.41 | 47.89 |
| 0.6447 | 0.5002 | 1105 | 0.6348 | 49.23 | 49.23 | 56.87 |