Views
No views yet
EleutherAI/pythia-70m instead. This repository is intended for research that requires access to optimizer states or the original training format.layer_XX-model_00-model_states.pt — model weight shards (one per layer)mp_rank_00_model_states.pt — model state metadatazero_pp_rank_*_optim_states.pt — ZeRO optimizer states (Adam moments, etc.)70M.yml — GPT-NeoX training configurationstep0 — initializationstep{1,2,4,8,16,32,64,128,256,512} — log-spaced early checkpointsstep1000 through step143000 — every 1,000 stepsstep143000 corresponds to the final model.1python tools/convert_neox_to_hf.py \
2 --input_dir /path/to/neox/checkpoint \
3 --config_file /path/to/config.yml \
4 --output_dir /path/to/hf/outputEleutherAI/pythia-70m.| Pythia Model | Non-Embedding Params | Layers | Model Dim | Heads | Batch Size | Learning Rate |
|---|---|---|---|---|---|---|
| 70M | 18,915,328 | 6 | 512 | 8 | 2M | 1.0 x 10-3 |
| 160M | 85,056,000 | 12 | 768 | 12 | 2M | 6.0 x 10-4 |
| 410M | 302,311,424 | 24 | 1024 | 16 | 2M | 3.0 x 10-4 |
| 1B | 805,736,448 | 16 | 2048 | 8 | 2M | 3.0 x 10-4 |
| 1.4B | 1,208,602,624 | 24 | 2048 | 16 | 2M | 2.0 x 10-4 |
| 2.8B | 2,517,652,480 | 32 | 2560 | 32 | 2M | 1.6 x 10-4 |
| 6.9B | 6,444,163,072 | 32 | 4096 | 32 | 2M | 1.2 x 10-4 |
| 12B | 11,327,027,200 | 36 | 5120 | 40 | 2M | 1.2 x 10-4 |
1@article{biderman2023pythia,
2 title={Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling},
3 author={Biderman, Stella and Schoelkopf, Hailey and Anthony, Quentin Gregory and Bradley, Herbie and O'Brien, Kyle and Hallahan, Eric and Khan, Mohammad Aflah and Purohit, Shivanshu and Prashanth, USVSN Sai and Raff, Edward and others},
4 journal={International Conference on Machine Learning},
5 year={2023}
6}