Views
No views yet
output_router_logits=True1Architecture:
2 - Layers: 48
3 - Hidden Size: 2048
4 - Attention Heads: 32
5 - Experts per Layer: 90 (reduced from 128)
6 - Active Experts per Token: 8
7 - Context Length: 128K
8 - Effective Parameters: ~21B (reduced from ~30B)
9
10Optimizations:
11 - FP8 quantization preserved
12 - SafeTensors format
13 - Flash Attention compatible
14 - Efficient expert routing
15 - True architectural pruning
16
17
18| Metric | Original | Pruned | Change |
19|--------|----------|--------|--------|
20| Total Parameters | ~30B | ~21B | -28.2% |
21| Model Size | 56.9 GB | 40.8 GB | -16.0 GB |
22| Experts per Layer | 128 | 90 | -38 |
23| Evaluation Loss | Baseline | +7.36% | Probable degradation |
24