Views
No views yet
qwen2.5-0.5b121a, scalar_flash, legacy_bf16, optimizer lionhf_source_checkpoint/ — original Hugging Face checkpoint and tokenizer used as the starting point.native_checkpoint/ — checkpoint converted to the IDA native format before training.native_input/ — packed SQuAD token/label/segment binary blocks consumed by the native trainer.run/ — full training outputs: status, metrics, logs, final native checkpoint, optimizer state, and (for GPT-2) torch fallback checkpoint.| field | value |
|---|---|
| architecture | ida_lattice |
| hidden size | 896 |
| layers | 24 |
| heads | 14 |
| intermediate size | 4864 |
| vocab size | 151936 |
| metric | value |
|---|---|
| phase | native_smoke_complete |
| final loss | 8.685278 |
| wallclock (s) | 2226.2875 |
| micro-steps | 2113 |
| optimizer steps | 529 |
| examples seen | 8452 |
| tokens seen | 17309696 |
| sequence length | 2048 |
| microbatch | 4 |
| grad accumulation | 4 |
| global batch size | 16 |
| base learning rate | 0.0003 |
| throughput | 7775.14 tok/s |
| steps/sec | 0.949 |
| ms/step | 1053.61 |
| optimizer steps/sec | 0.238 |
| skipped optimizer steps | 415 |
| final grad norm | 112485984.0 |
| approx params | 630,167,424 |
| compute capability | 12.1 |
| attention backend | scalar_flash |
| precision profile | legacy_bf16 |
KissTheHabit/IDA-TRAIN-V2-mlperf-squad-gb10-results.