Views
No views yet
[!IMPORTANT] This repository is an experimental quantized version of the original modelibm-granite/granite-3.1-3b-a800m-base.It requires development versions oftransformersandbitsandbytes.
nn.Linear modules except lm_head and router modules, using an experimental bnb_4bit_target_parameters configuration option.| Model | 2B Dense | 8B Dense | 1B MoE | 3B MoE |
|---|---|---|---|---|
| Embedding size | 2048 | 4096 | 1024 | 1536 |
| Number of layers | 40 | 40 | 24 | 32 |
| Attention head size | 64 | 128 | 64 | 64 |
| Number of attention heads | 32 | 32 | 16 | 24 |
| Number of KV heads | 8 | 8 | 8 | 8 |
| MLP hidden size | 8192 | 12800 | 512 | 512 |
| MLP activation | SwiGLU | SwiGLU | SwiGLU | SwiGLU |
| Number of Experts | — | — | 32 | 40 |
| MoE TopK | — | — | 8 | 8 |
| Initialization std | 0.1 | 0.1 | 0.1 | 0.1 |
| Sequence Length | 4096 | 4096 | 4096 | 4096 |
| Position Embedding | RoPE | RoPE | RoPE | RoPE |
| # Parameters | 2.5B | 8.1B | 1.3B | 3.3B |
| # Active Parameters | 2.5B | 8.1B | 400M | 800M |
| # Training tokens | 12T | 12T | 10T | 10T |