Author / contact: @Kyokopom on X Repository: KitsuVp/NeoLLM
| Parameter | Value |
|---|---|
| Hidden size | 512 |
| Layers | 12 |
| Attention heads | 8 |
| KV heads (GQA) | 4 |
| Head dim | 64 |
| Intermediate size | 1536 |
| Vocabulary | LiquidAI/LFM2.5-1.2B-Thinking tokenizer (64,402 tokens) |
| Context length | 512 tokens |
| Parameter bucket | Count |
|---|---|
| Total parameters | 84.57M (84,567,928) |
| Embedding parameters (tied) | 32.97M (32,973,824) |
| Non-embedding parameters | 51.59M (51,594,104) |
| Effective trainable parameters | 84.57M (84,567,928) |
Weight tying is enabled: the input embedding matrix and the language-model head share the same parameters, so the effective trainable budget istotal − embed = 51.59M.
[α·softmax(QKᵀ) + β]·V.lambda * ||mean(output_embeddings)||^2 regularizer for output-logit stability.| Feature | Enabled | Value |
|---|---|---|
| MiLe Loss | True | gamma=1.0 |
| mu-loss | True | lambda=0.0001 |
| MEAP | True | ratio=0.15 |
Kitsunp/ml-cross-entropy package only when
their corresponding flags are enabled. With all flags disabled, NeoLLM calls upstream CCE
without extension-specific arguments. When any extension is active, CCE reports three compact
scalars: unweighted NTP cross entropy, the MiLe reweighting delta, and the mu-loss penalty.
Their sum reconstructs ntp_loss exactly. MEAP reports its eligible and selected counts from
the masking kernel; the trainer logs the selected count and exact fraction. No diagnostic path
materializes full-vocabulary logits or a token mask outside the kernels.ADEMAMIX=True. DeltaMomentum
(arXiv:2608.19491) optionally replaces only the fast beta1
first-moment update on hidden Linear weights when DELTA=True; beta2 remains Adam's
ordinary squared-gradient EMA in every mode.| Setting | Value |
|---|---|
| Dataset | FineWeb-Edu (sample-10BT) |
| Tokens seen | ~0.51B (15,625 steps × batch 64 × length 512) |
| Precision | BF16 model state (bfloat16: 84,565,816 elements, float32: 2,112 elements) + optimizer policy unavailable |
| Optimizer | AdEMAMix |
| Gradient transform | Active; repository per-parameter tensor norms; bias correction=True; no stabilizer warmup; EMA state stored in optimizer checkpoints |
| Learning rate | 1e-03 with 4,000-step linear warmup |
| Weight decay | coefficients [0.005, 0.01, 0.1]; WD 5/5 groups; CWD 2; Huber 0; Correction 0 |
| Training time | ~1h 08m |
| Hardware | NVIDIA GeForce RTX 5090 |
| Step | Train Loss | Val Loss |
|---|---|---|
| 5,000 | 5.641 | 5.236 |
| 10,000 | 5.374 | 4.921 |
| 15,000 | 5.204 | 4.745 |
| 15,625 | — | 4.726 |
| Area | Technique | Paper title | Reference |
|---|---|---|---|
| Embeddings | Learnable Multipliers | Freeing the Scale of Language Model Matrix Layers | arXiv:2601.04890 |
| Embeddings | Leviathan | A Separable Architecture for Continuous Token Representation in Language Models | arXiv:2601.22040 |
| Embeddings | KHRONOS | KHRONOS: a Kernel-Based Neural Architecture for Rapid, Resource-Efficient Scientific Computation | arXiv:2505.13315 |
| Embeddings | Spelling Bee | Spelling Bee Embeddings for Language Modeling | arXiv:2601.18030 |
| Embeddings | Token embedding analysis | Token Embeddings Violate the Manifold Hypothesis | arXiv:2504.01002 |
| Attention / positions | FAN | Fourier Analysis Networks | arXiv:2502.21309 |
| Attention / positions | MEA | Explicit Multi-head Attention for Inter-head Interaction in Large Language Models | arXiv:2601.19611 |
| Attention / positions | LUCID | Attention with Preconditioned Representations | arXiv:2602.10410 |
| Attention / positions | Affine-Scaled Attention | Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention | arXiv:2602.23057 |
| Attention / positions | XSA | Exclusive Self Attention | arXiv:2603.09078 |
| Attention / positions | Directional Routing | Directional Routing in Transformers | arXiv:2603.14923 |
| Attention / positions | Gated Attention | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free | arXiv:2505.06708 |
| Attention / positions | Momentum Attention | Momentum Attention | arXiv:2411.03884 |
| Attention / positions | IHA | Interleaved Head Attention | arXiv:2602.21371 |
| Attention / positions | REPO | Language Models with Context Re-Positioning | arXiv:2512.14391 |
| Attention / positions | GRAPE | Group Representational Position Encoding | arXiv:2512.07805 |
| Attention / positions | GOAT priors | You Need Better Attention Priors | arXiv:2601.15380 |
| Attention / positions | Hadamard o_proj | Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers | arXiv:2603.08343 |
| Residual / normalization | SeeDNorm | Self-Rescaled Dynamic Normalization | arXiv:2510.22777 |
| Residual / normalization | LNS | The Curse of Depth in LLMs | arXiv:2502.05795 |
| Residual / normalization | GPAS | Gradient-Preserving Activation Scaling | arXiv:2506.22049 |
| Residual / normalization | PolyNorm | PolyNorm / PolyCom | arXiv:2602.04902 |
| Residual / normalization | SimpleGPT | SimpleGPT | arXiv:2602.01212 |
| Residual / normalization | StackMemory / STACKTRANS | Recursive Transformer: Boosting Reasoning Ability with State Stack | NeurIPS 2025 |
| Residual / normalization | Attention Residuals | Attention Residuals | arXiv:2603.15031 |
| Residual / normalization | LAUREL | LAUREL: Learned Augmented Residual Layer | arXiv:2411.07501 |
| Objectives | TWEO | Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies | arXiv:2511.23225 |
| Objectives | NITP | Next Implicit Token Prediction for LLM Pre-training | arXiv:2605.24956 |
| Objectives | NextLat | Next-Latent Prediction Transformers Learn Compact World Models | arXiv:2511.05963 |
| Optimizer / training | AdEMAMix | The AdEMAMix Optimizer: Better, Faster, Older | arXiv:2409.03137 |
| Optimizer / training | DeltaMomentum | DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule | arXiv:2608.19491 |
| Optimizer / training | BF16 + SR | Stochastic Rounding for LLM Training: Theory and Practice | arXiv:2502.20566 |
| Optimizer / training | CWD | Cautious Weight Decay | arXiv:2510.12402 |
| Optimizer / training | WD correction | Correction of Decoupled Weight Decay | arXiv:2512.08217 |
| Optimizer / training | AdamHD | AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training | arXiv:2511.14721 |
| Optimizer / training | GradientStabilizer | GradientStabilizer | arXiv:2502.17055 |
1@misc{neollm2026,
2 title = {NeoLLM: A Research Language Model Integrating Recent Attention and Normalization Techniques},
3 author = {KitsuVp},
4 year = {2026},
5 url = {https://huggingface.co/KitsuVp/NeoLLM}
6}