Views
No views yet
| Component | T4NT-0.5B |
|---|---|
| Parameters | 535,033,600 (0.54B) |
| Embedding dim | 1280 |
| Attention heads | 16 |
| Layers | 20 |
| FFN dim | 3584 |
| Max sequence length | 256 |
| Position encoding | Rotary (RoPE) |
| FFN activation | SwiGLU |
| Normalization | RMSNorm |
| Weight quantization | 4-bit symmetric (15 levels) |
| Weight clipping | tanh(w/3) * 3 |
| Vocab size | 50,257 (GPT-2 tokenizer) |
| Detail | Value |
|---|---|
| Dataset | WikiText-103 |
| Tokens seen | ~15M |
| Epochs | 50 |
| Batch size | 2 (effective 32 with gradient accumulation) |
| Learning rate | 3e-4 (cosine annealing) |
| Optimizer | AdamW (betas 0.9, 0.95) |
| Mixed precision | FP16 autocast + GradScaler |
| Hardware | NVIDIA T4 15GB |
| Training time | 2.92 hours |
| FP32 size | 2.04 GB |
| INT4 deployment size | 255 MB |
| Epoch | Train Loss | Val Loss | PPL |
|---|---|---|---|
| 1 | 7.985 | 7.254 | 1413.75 |
| 5 | 6.116 | 5.931 | 376.41 |
| 10 | 5.630 | 5.508 | 246.76 |
| 15 | 5.365 | 5.280 | 196.34 |
| 20 | 5.182 | 4.995 | 147.72 |
| 25 | 5.035 | 4.848 | 127.47 |
| 30 | 4.879 | 4.760 | 116.80 |
| 35 | 4.781 | 4.672 | 106.86 |
| 40 | 4.760 | 4.612 | 100.73 |
| 45 | 4.716 | 4.574 | 96.91 |
| 50 | 4.726 | 4.610 | 100.46 |
| Best | - | 4.533 | 93.04 |
| Metric | Epoch 1 | Epoch 25 | Epoch 50 |
|---|---|---|---|
| Weight std | 0.0201 | 0.0205 | 0.0199 |
| Weight abs_max | 0.1183 | 0.1607 | 0.1130 |
| Quantization levels used | 14.8 / 15 | 14.9 / 15 | 15.0 / 15 |
| Weights in [-3, 3] | 100% | 100% | 100% |
| Gradient norm | 2.22 | 0.61 | 0.60 |
The meaning of life is unclear, but it is still a significant, mature, obvious, a man of other males who would find the way of the female. The man died on April 8, 1820, and the father of the family, John, died on March 2
In the history of the first two years of the period. In his second career in 1848, Mary was named for his junior year by the University of North America, and later, in 1875, was named after his mother, William, a member of the University
1import torch
2
3# Load checkpoint
4ckpt = torch.load('T4NT_0.5B_tanh_50ep.pt', map_location='cpu')
5print(f"Val Loss: {ckpt['val_loss']:.4f}")
6print(f"PPL: {ckpt['ppl']:.2f}")
7print(f"Config: {ckpt['config']}")
8
9# Rebuild model architecture and load weights
10# See config.json for architecture parameters
11# Full training code available in the training logs| File | Description |
|---|---|
config.json | Model architecture and training configuration |
generation_config.json | Default generation parameters |
T4NT_0.5B_tanh_50ep.pt | Model checkpoint (FP32 training weights, 2.14 GB) |
t4nt_tanh_50ep_log.json | Detailed training logs with per-epoch metrics and weight statistics |
1@misc{tathe2025t4nt,
2 author = {Tathe, Shivnath},
3 title = {T4NT-0.5B: Tanh 4-bit Neural Transformer},
4 year = {2025},
5 url = {https://huggingface.co/shivnathtathe/T4NT-0.5B}
6}
7
8@misc{tathe2026true4bit,
9 author = {Tathe, Shivnath},
10 title = {True 4-Bit Quantized Convolutional Neural Network Training on CPU: Achieving Full-Precision Parity},
11 year = {2026},
12 eprint = {2603.13931},
13 archivePrefix = {arXiv},
14 primaryClass = {cs.LG}
15}