Views
No views yet
0.9.11base_model: meta-llama/Llama-3.1-8B-Instruct
2model_type: LlamaForCausalLM
3tokenizer_type: AutoTokenizer
4gradient_accumulation_steps: 2
5micro_batch_size: 4
6num_epochs: 1
7optimizer: adamw_bnb_8bit
8lr_scheduler: cosine
9learning_rate: 0.0001
10load_in_8bit: true
11load_in_4bit: false
12adapter: lora
13lora_model_dir: null
14lora_r: 8
15lora_alpha: 16
16lora_dropout: 0.05
17lora_target_modules:
18- q_proj
19- v_proj
20- k_proj
21datasets:
22- path: /workspace/FinLoRA/data/train/xbrl_term_train.jsonl
23 type:
24 system_prompt: ''
25 field_system: system
26 field_instruction: context
27 field_output: target
28 format: '[INST] {instruction} [/INST]'
29 no_input_format: '[INST] {instruction} [/INST]'
30dataset_prepared_path: null
31val_set_size: 0.02
32output_dir: /workspace/FinLoRA/lora/axolotl-output/xbrl_term_llama_3_1_8b_8bits_r8
33peft_use_dora: false
34sequence_len: 4096
35sample_packing: false
36pad_to_sequence_len: false
37wandb_project: finlora_models
38wandb_entity: null
39wandb_watch: gradients
40wandb_name: xbrl_term_llama_3_1_8b_8bits_r8
41wandb_log_model: 'false'
42bf16: auto
43tf32: false
44gradient_checkpointing: true
45resume_from_checkpoint: null
46logging_steps: 500
47flash_attention: false
48deepspeed: deepspeed_configs/zero1.json
49warmup_steps: 10
50evals_per_epoch: 4
51saves_per_epoch: 1
52weight_decay: 0.0
53special_tokens:
54 pad_token: <|end_of_text|>
55chat_template: llama3
56| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| No log | 0.0070 | 1 | 2.5692 |
| No log | 0.2509 | 36 | 1.7055 |
| No log | 0.5017 | 72 | 1.5480 |
| No log | 0.7526 | 108 | 1.5077 |