Views
No views yet
0.4.11base_model: Qwen/Qwen2-0.5B
2batch_size: 32
3bf16: true
4chat_template: tokenizer_default_fallback_alpaca
5datasets:
6- data_files:
7 - 745d2d05aaed18f4_train_data.json
8 ds_type: json
9 format: custom
10 path: /workspace/input_data/745d2d05aaed18f4_train_data.json
11 type:
12 field_input: pos
13 field_instruction: task
14 field_output: query
15 format: '{instruction} {input}'
16 no_input_format: '{instruction}'
17 system_format: '{system}'
18 system_prompt: ''
19eval_steps: 20
20flash_attention: true
21gpu_memory_limit: 80GiB
22gradient_checkpointing: true
23group_by_length: true
24hub_model_id: willtensora/459779f2-cbce-4ec0-b11c-1dcdf92498d8
25hub_strategy: checkpoint
26learning_rate: 0.0002
27logging_steps: 10
28lr_scheduler: cosine
29max_steps: 2500
30micro_batch_size: 4
31model_type: AutoModelForCausalLM
32optimizer: adamw_bnb_8bit
33output_dir: /workspace/axolotl/configs
34pad_to_sequence_len: true
35resize_token_embeddings_to_32x: false
36sample_packing: false
37save_steps: 40
38save_total_limit: 1
39sequence_len: 2048
40tokenizer_type: Qwen2TokenizerFast
41train_on_inputs: false
42trust_remote_code: true
43val_set_size: 0.1
44wandb_entity: ''
45wandb_mode: online
46wandb_name: Qwen/Qwen2-0.5B-/workspace/input_data/745d2d05aaed18f4_train_data.json
47wandb_project: Gradients-On-Demand
48wandb_run: your_name
49wandb_runid: default
50warmup_ratio: 0.05
51xformers_attention: true
52| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| No log | 0.0004 | 1 | 3.9660 |
| 2.8207 | 0.0086 | 20 | 3.1038 |
| 3.1247 | 0.0172 | 40 | 3.0989 |
| 2.9411 | 0.0258 | 60 | 2.8986 |
| 2.9915 | 0.0344 | 80 | 2.8742 |
| 2.8038 | 0.0430 | 100 | 2.8405 |
| 2.8518 | 0.0516 | 120 | 2.7728 |
| 2.7079 | 0.0602 | 140 | 2.6985 |
| 2.6076 | 0.0688 | 160 | 2.6416 |
| 2.6172 | 0.0774 | 180 | 2.5695 |
| 2.552 | 0.0860 | 200 | 2.5151 |
| 2.5036 | 0.0946 | 220 | 2.4783 |
| 2.4887 | 0.1032 | 240 | 2.4610 |
| 2.4008 | 0.1118 | 260 | 2.4569 |
| 2.424 | 0.1204 | 280 | 2.4560 |