FrontD.11m
FrontD.11m is an approximately 11-million-parameter decoder-only
language model trained completely from scratch.
Architecture
- Parameters: 10,996,608
- Context: 1096
- Vocabulary: 16000
- Hidden size: 256
- Layers: 6
- Attention heads: 16
- FFN size: 1152
- RoPE
- RMSNorm
- SwiGLU
- Tied input/output embeddings
Pretraining dataset
Plans11/Organized_PreTrain_1k_Context
Training
Training is performed in 20,000-example sessions.
A persistent SHA-256 example ledger records only examples whose
training session has successfully completed.
The model is initialized randomly and pretrained from scratch.