Views
No views yet
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 6.3991 | 1.0 | 2249 | 6.5169 |
| 6.2508 | 2.0 | 4498 | 6.3377 |
gpt2 model checkpoint.wikitext-2-raw-v1 dataset.2e-50.01per_device_train_batch_size in TrainingArguments (default is 8). Gradient accumulation steps are not set (default is 1).Trainer)Trainer)