Dense-306
Dense-306 is the 306.5M-parameter dense baseline from the matched run pair in
Token Identity as a Routing Signal for Residual MLP Experts. It is a
decoder-only base language model trained from scratch on approximately 8B
FineWeb-Edu tokens.
This is a pretrained base model, not an instruction-tuned or safety-aligned
assistant.
Architecture
- 18 transformer layers
- hidden size 1,024
- 16 query heads and 4 key/value heads (GQA)
- QK-Norm and RoPE
- 2,048-token training context
- 32,000-token tokenizer
- dense SwiGLU width 4,096
- 306,486,528 trainable parameters
Evaluation
Both models were evaluated zero-shot with EleutherAI LM Evaluation Harness
0.4.12, identical seeds, BF16 inference, no sample limit, and a maximum context
of 2,048 tokens.
| Model | ARC-Easy acc_norm | PIQA acc_norm | HellaSwag acc_norm | WikiText-2 word PPL |
|---|
| Dense-306 | 49.07 | 63.49 | 33.67 | 35.79 |
| TR-MOE-306 | 48.61 | 63.49 | 34.06 | 35.20 |
The multiple-choice differences are smaller than one reported standard error.
These results support near-parity, not task-level superiority.
At the last common point on a fixed FineWeb-Edu training-split evaluation
stream, the dense model reached NLL 2.9482 and the token-routed model reached
2.9329. That stream is not an independent held-out set, and each architecture
was trained with one seed.
Files and integrity
model.safetensors is the exact FP32 model export used for the reported
downstream evaluation.
- final training step:
7629
- training checkpoint SHA-256:
0d5bc1d8d8ead8ecde70e2705dfc8712e99a3de79bdf24f9b4e34fd1f668de7b
- exported
model.safetensors SHA-256:
311c97b2233a0d182fe47fa641a427e30f59b49774872a85b59854457ed1e7ae
The repository also contains the tokenizer, generation configuration, and
export_manifest.json.
Runtime
The checkpoint uses the custom
DeepForCausalLM format implemented by the
Complexity-ML vLLM runtime.
Stock Transformers and upstream vLLM do not currently register the
deep
architecture. Use that runtime for inference, or adapt its loader before
loading these weights.
The tokenizer can be loaded independently:
1from transformers import AutoTokenizer
2
3tokenizer = AutoTokenizer.from_pretrained("Pacific-i64/Dense-306")
Limitations
This is a small research base model. It may generate inaccurate, biased,
offensive, repetitive, or unsafe text. It should not be used for high-stakes
decisions. The training corpus is web-derived and may contain errors,
copyrighted material, personal information, and social biases.
Paper and reproducibility
The paper reports a matched single-seed observation. Multi-seed replication
remains open.
License
The Dense-306 model weights are released under the Apache License 2.0. See
LICENSE.