The root model.safetensors is the epoch 3 / step
5,982 checkpoint, selected for the strongest full-PIQA result
and lowest held-out SFT loss across the three epochs.
Tokenizer compatibility notice (23 August 2026): this released SFT
remains a 32,000-token checkpoint and must be loaded with the tokenizer files
stored in this repository. It was trained from the pre-migration 32,000-token
Refinement revision. The canonical Base and Refinement roots now use 32,004
tokens; a replacement SFT must be retokenized from raw text and explicitly
train IDs 32,000–32,003. Do not pair this checkpoint with the standalone
32,004-token tokenizer or resize it as a substitute for retraining.
Results
Epoch
Step
Held-out SFT loss
SFT ppl
PIQA acc
PIQA acc_norm
1
1,994
0.990943
2.69
67.90%
68.93%
2
3,988
0.963912
2.62
67.85%
68.82%
3
5,982
0.959617
2.61
68.01%
69.10%
Released checkpoint benchmark panel
Benchmark
Split
Examples
Accuracy
Accuracy (length-normalized)
Evaluation backend
PIQA
validation
1,838
68.01%
69.10%
PyTorch FP16, custom Triton
ARC-Easy
test
2,376
57.24%
52.86%
MLX FP16
ARC-Challenge
test
1,172
27.13%
29.01%
MLX FP16
ARC Combined (micro)
test
3,548
47.29%
44.98%
Derived from both ARC test splits
HellaSwag
validation
10,042
33.21%
38.74%
MLX FP16
All benchmark evaluations use zero-shot causal continuation log-likelihood, no
chat template and a maximum sequence length of 2,048. Accuracy selects the
choice with the highest total continuation log-likelihood; the normalized
metric selects by mean continuation log-likelihood per scored token. PIQA was
evaluated from the native epoch-3 checkpoint. ARC and HellaSwag were
evaluated from an FP16 MLX conversion of the same root F32 SafeTensors weights.
ARC Combined is the micro-average over all 3,548 ARC-Easy and ARC-Challenge
test examples, not the arithmetic mean of the two percentages.
Machine-readable reports are published under reports/sft-v2-300k/.
Compact-model comparison
Combined ARC comparison from 124M to 774M parameters
Model
Parameters
PIQA acc
ARC-Easy acc
ARC-Challenge acc
Combined ARC acc
HellaSwag acc
TR-HASH MoE 200M Full SFT
201.2M
68.01%
57.24%
27.13%
47.29%
33.21%
GPT-2 Large
774M
—
53.11%
21.76%
42.76%
—
Pythia-410M
410M
—
52.02%
21.42%
41.91%
—
GPT-2 Medium
355M
—
49.16%
21.67%
40.08%
—
OPT-350M
350M
—
43.98%
20.82%
36.33%
—
GPT-2 Small
124M
62.89%
43.81%
19.03%
≈35.63%
28.92%
OPT-125M
125M
63.00%
43.60%
19.10%
≈35.51%
29.20%
Pythia-160M
160M
62.73%
43.52%
18.77%
35.34%
—
Combined ARC is weighted by the public test-set sizes (2,376 ARC-Easy and
1,172 ARC-Challenge examples). GPT-2 Small and OPT reference scores come from
the AMD-LLM lm-evaluation-harness comparison;
their combined values are approximate because the published component scores
are rounded. Pythia-160M uses EleutherAI's
official zero-shot result.
GPT-2 Medium, GPT-2 Large and Pythia-410M were evaluated locally in MLX FP16
with the same causal-choice evaluator as TR-HASH. OPT-350M used the identical
prompt and scoring formula in PyTorch MPS FP16 because MLX does not implement
the OPT architecture. The 124M–160M references were reported through
lm-evaluation-harness, so that part of the comparison is informative rather
than a claim of bit-identical evaluation runtimes.
Final assistant turn only; prior assistant turns masked
Epochs
3
Context
2,048 tokens
Tokenizer
TR-HASH 32,000-token vocabulary; EOS </s> (ID 0)
Optimizer
AdamW, betas 0.9 / 0.95, weight decay 0.1
LR
2e-5 peak, 3% warmup, continuous cosine decay
Precision
BF16 training
Root SafeTensors precision
float32
Kernels
Liger required; custom Triton enabled
Architecture and loading
201.2M parameters, 16 decoder layers, GQA (14 query heads / 2 KV heads),
four stored deterministic token-ID-routed experts with top-2 activation, an
always-on shared SwiGLU path and tied embeddings. The persisted multi-hash
routing tables are part of the checkpoint.
The repository includes an autonomous Transformers adapter. Load it with: