MoE-Study — Dense vs. Mixture-of-Experts, matched active parameters
Two decoder-only language models trained from scratch under identical conditions, differing in
exactly one thing: whether the feed-forward block is a dense MLP or a sparse top-2-of-4 MoE.
The MoE's active-parameter count matches Dense by construction — 2 of 4 experts at half the hidden
size means identical compute per token. The MoE only spends more memory for extra capacity.
⚠️ These are research artifacts, not usable models
Read this before downloading.
Trained for one epoch on ~40.7M tokens — neither model is close to converged.
WikiText word perplexity is 551 (Dense) and 1,378 (MoE). Generations are largely incoherent.
0.0% on LAMBADA for both — at the task floor.
No instruction tuning, no RLHF, no safety filtering of any kind.
They exist to answer one narrow question: at matched active compute and matched budget, does sparsity
help? They are not fit for any downstream use.
Getting the weights
These are a custom architecture, not a variant of an existing one. The modeling code is not included
here, so from_pretrained on this repo alone will not build the model.
Clone the GitHub repo — it carries the model definition
and loading instructions, and points back at these subfolders for the weights.
Model details
Shared architecture
Both models are the same custom decoder-only transformer:
Layers
12
Attention heads
12
Embedding dim
768
Context length
1024
Vocabulary
50,257 (GPT-2 tokenizer)
Attention
Multi-Query — one shared K/V projection across all query heads
Normalization
Custom pre-norm (learned scale + shift)
Position embeddings
Learned absolute
Weight tying
None — separate input embedding and output head
The one difference
dense/
moe/
FFN block
2-layer GELU MLP
4 experts, top-2 routed
hidden_dim
3072
1536 (per expert)
Router
—
linear → softmax → top-2, renormalized
Aux loss
—
load-balancing term, summed over all 12 layers
Both models share the same unmodified GPT-2 tokenizer, stored once at the repo root.
Training
Identical for both models. Single consumer GPU, no cloud.
Dense wins on everything sensitive to raw LLM quality — perplexity, ARC-Easy, PIQA.
WinoGrande, HellaSwag, and ARC-Challenge are won by MoE, but with such a negligible difference, that they could be considered to have an equal accuracy
Inference speed
Greedy decoding, 32-token prompt → 64 new tokens, 5 trials, 2 warmup, no KV cache.
Model
Tokens/sec
Total params
Active params/token
Dense
106.49 ± 0.30
150.1M
150.1M
MoE
34.40 ± 0.08
206.8M
~150.1M
MoE is ~3.1× slower despite matched active compute — an artifact of unoptimized expert dispatch, not
a property of the architecture.
Benchmark charts
ARC-Easy
PIQA
WikiText
LAMBADA
WinoGrande
HellaSwag
ARC-Challenge
Speed
Findings
1. Dense won every metric that wasn't already at chance.
Most clearly on WikiText perplexity — 551 vs 1,378, a 2.5× gap.
2. The routing math is correct.
Active parameters match Dense almost exactly. Matched active compute simply didn't buy matched quality
at this budget.
3. Routing stayed balanced.
The load-balancing term sat on its theoretical floor, so the MoE's gap is not explained by experts
collapsing onto each other.
4. Extra capacity needs extra tokens.
The MoE has 38% more parameters but saw the same ~40.7M tokens — likely far too few to train 4 experts
per layer, each seeing only a routed fraction of the stream.
Citation
bibtex
1@misc{sundaram2026moestudy,
2 author = {Sundaram, Oliver},
3 title = {MoE-Study: Dense vs. Mixture-of-Experts at Matched Active Parameters},
4 year = {2026},
5 url = {https://github.com/OliverSundaram/MoE-Study}
6}