Views
No views yet
0.10.0.dev01adapter: lora
2base_model: Qwen/Qwen2.5-0.5B
3bf16: true
4chat_template: llama3
5datasets:
6- data_files:
7 - 4803e74546490252_train_data.json
8 ds_type: json
9 format: custom
10 path: /workspace/input_data/
11 type:
12 field_input: None
13 field_instruction: french
14 field_output: wolof
15 field_system: None
16 format: None
17 no_input_format: None
18 system_format: '{system}'
19 system_prompt: ''
20eval_max_new_tokens: 256
21evals_per_epoch: 2
22flash_attention: false
23fp16: false
24gradient_accumulation_steps: 4
25gradient_checkpointing: true
26group_by_length: true
27hub_model_id: apriasmoro/59c50be6-3cd6-4959-9504-237de1f26d7d
28learning_rate: 0.0002
29logging_steps: 10
30lora_alpha: 16
31lora_dropout: 0.05
32lora_fan_in_fan_out: false
33lora_r: 8
34lora_target_linear: true
35lr_scheduler: cosine
36max_steps: 10
37micro_batch_size: 16
38mlflow_experiment_name: /tmp/4803e74546490252_train_data.json
39model_type: AutoModelForCausalLM
40num_epochs: 3
41optimizer: adamw_bnb_8bit
42output_dir: miner_id_24
43pad_to_sequence_len: true
44sample_packing: false
45save_steps: 200
46sequence_len: 2048
47tf32: true
48tokenizer_type: AutoTokenizer
49train_on_inputs: false
50trust_remote_code: true
51val_set_size: 0.05
52wandb_entity: null
53wandb_mode: online
54wandb_name: ac83fef5-7c82-4619-8ae0-c818a659c38b
55wandb_project: Gradients-On-Demand
56wandb_run: apriasmoro
57wandb_runid: ac83fef5-7c82-4619-8ae0-c818a659c38b
58warmup_steps: 100
59weight_decay: 0.01
60| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| No log | 0.0041 | 1 | 7.8505 |
| No log | 0.0082 | 2 | 7.8608 |
| No log | 0.0163 | 4 | 7.8459 |
| No log | 0.0245 | 6 | 7.8270 |
| No log | 0.0327 | 8 | 7.8066 |
| 6.8006 | 0.0409 | 10 | 7.7619 |