Shivik-1B-MEGA
Description
Ultimate merged model combining 2 unique Shivik-1B variants.
Merged Models
Architecture
- Parameters: ~1.24B
- Hidden size: 2048
- Layers: 16
- Attention heads: 32
- KV heads: 8 (GQA 4:1)
- Vocab size: 128,262
Special Tokens
<think> / </think> - Chain-of-thought reasoning
<step> / </step> - Step-by-step problem solving
<answer> / </answer> - Final answers
Note
Models excluded from merge (incompatible architecture):
- Teacher-Enhanced: 1.24B, norm=torch.Size([3072])
- Trained-Merged: 1.50B, norm=torch.Size([2048])