Zenyx V3 Base (1.5B Mixture-of-Experts)
Zenyx V3 is an efficient 1.5B-parameter Mixture-of-Experts (MoE) foundation model
built for low-latency inference and high throughput. It is written from scratch in
JAX/Flax and trained on TPU v5e-8.
[!IMPORTANT]
This is a BASE model — it is not instruction-tuned. It completes text; it does
not follow instructions or hold a conversation. Prompt it with a prefix to continue
("The capital of France is"), not with a request ("Explain gravity").
Pretraining is still in progress; SFT/chat variants will follow.
Current checkpoint: step 86,400 · 59.1B tokens seen
Model Architecture
- Sparse Mixture-of-Experts: 12 routed experts + 1 shared expert, exactly 2
active per token, with a Sinkhorn transport-based gate.
- Multi-head Latent Attention (MLA): compresses the KV cache into a low-rank
latent subspace, cutting HBM bandwidth and memory footprint.
- Hyper-Connections: Sinkhorn-normalised residual routing for gradient
stability at scale.
- Multi-Token Prediction (MTP): one auxiliary prediction head during training.
- Context: pretrained at up to 4,096 tokens (progressive 2,048 → 4,096).
YaRN and RoPE scaling factors are precomputed so context can be extended at
inference time beyond the trained length.
| |
|---|
| Total parameters | ~1.5B |
| Active parameters / token | ~0.4B |
| Layers | 16 (2 dense + 14 MoE) |
| Hidden size | 1,536 |
| Attention heads | 12 (head dim 128) |
| Vocabulary | 129,280 |
| Precision | bfloat16 |
Benchmarks — checkpoint step 86,400 (59.1B tokens)
All tasks are evaluated with the standard base-model protocol: the model scores
the log-likelihood of every candidate continuation and the highest-scoring one is
taken as the answer. Nothing is generated and no output parsing is involved, so the
numbers do not depend on instruction-following ability. 0-shot, full evaluation
sets, no subsampling.
acc_norm normalises each continuation's log-likelihood by its length in characters,
which removes the bias toward short answers; it is the headline metric wherever the
task has candidates of differing lengths.
| Benchmark | acc | acc_norm | Random | Δ | n | Description |
|---|
| HellaSwag | 30.18% ± 0.46 | 33.67% ± 0.47 | 25.0% | +8.7 | 10,042 | Commonsense sentence completion |
| ARC-Easy | 50.63% ± 1.03 | 46.25% ± 1.02 | 25.0% | +21.3 | 2,376 | Grade-school science questions |
| ARC-Challenge | 21.93% ± 1.21 | 26.45% ± 1.29 | 25.0% | +1.5 | 1,172 | Hard grade-school science questions |
| PIQA | 60.72% ± 1.14 | 61.26% ± 1.14 | 50.0% | +11.3 | 1,838 | Physical commonsense reasoning |
| WinoGrande | 49.72% ± 1.40 | — | 50.0% | -0.3 | 1,267 | Pronoun resolution / coreference |
| OpenBookQA | 19.60% ± 1.78 | 31.00% ± 2.07 | 25.0% | +6.0 | 500 | Elementary science with open book |
| BoolQ | 60.76% ± 0.85 | 62.17% ± 0.85 | 62.2% | -1.4 | 3,270 | Yes/no reading comprehension |
| SciQ | 78.20% ± 1.31 | 69.80% ± 1.45 | 25.0% | +53.2 | 1,000 | Science exam questions with support |
| LAMBADA (OpenAI) | 25.64% ± 0.61 | — | 0.0% | +25.6 | 5,153 | Long-range last-word prediction |
| MMLU (5-shot) | 25.78% ± 0.37 | — | 25.0% | +0.8 | 14,042 | 57 subjects of academic knowledge |
| RACE | 30.34% ± 0.65 | 32.81% ± 0.67 | 25.0% | +7.8 | 4,934 | Exam reading comprehension |
| CommonsenseQA | 28.50% ± 1.29 | 31.04% ± 1.32 | 20.0% | +11.0 | 1,221 | 5-choice commonsense (random = 20%) |
| COPA | 62.00% ± 4.85 | 61.00% ± 4.88 | 50.0% | +12.0 | 100 | Causal reasoning |
| LogiQA | 21.51% ± 1.61 | 25.96% ± 1.72 | 25.0% | +1.0 | 651 | Logical deduction |
| WSC273 | 53.48% ± 3.02 | — | 50.0% | +3.5 | 273 | Winograd coreference |
| TruthfulQA MC1 | 19.58% ± 1.39 | — | 22.8% | -3.2 | 817 | Resistance to common misconceptions |
| Arithmetic | 0.09% ± 0.03 | — | 0.0% | +0.1 | 14,000 | 2-5 digit add/sub/mul, generated in-harness |
Bold marks the metric that is conventional for that task — acc_norm for
HellaSwag, ARC, PIQA and OpenBookQA; acc for WinoGrande, BoolQ, SciQ and
LAMBADA. The convention is applied per task, not chosen per result: it lowers
the reported figure for ARC-Easy (44.53 rather than 49.54) and PIQA (59.79
rather than 60.83). Δ compares the bolded metric to the random baseline.
Both metrics
| Benchmark | acc | acc_norm | n |
|---|
| HellaSwag | 30.18% ± 0.46 | 33.67% ± 0.47 | 10,042 |
| ARC-Easy | 50.63% ± 1.03 | 46.25% ± 1.02 | 2,376 |
| ARC-Challenge | 21.93% ± 1.21 | 26.45% ± 1.29 | 1,172 |
| PIQA | 60.72% ± 1.14 | 61.26% ± 1.14 | 1,838 |
| WinoGrande | 49.72% ± 1.40 | — | 1,267 |
| OpenBookQA | 19.60% ± 1.78 | 31.00% ± 2.07 | 500 |
| BoolQ | 60.76% ± 0.85 | 62.17% ± 0.85 | 3,270 |
| SciQ | 78.20% ± 1.31 | 69.80% ± 1.45 | 1,000 |
| LAMBADA (OpenAI) | 25.64% ± 0.61 | — | 5,153 |
| MMLU (5-shot) | 25.78% ± 0.37 | — | 14,042 |
| RACE | 30.34% ± 0.65 | 32.81% ± 0.67 | 4,934 |
| CommonsenseQA | 28.50% ± 1.29 | 31.04% ± 1.32 | 1,221 |
| COPA | 62.00% ± 4.85 | 61.00% ± 4.88 | 100 |
| LogiQA | 21.51% ± 1.61 | 25.96% ± 1.72 | 651 |
| WSC273 | 53.48% ± 3.02 | — | 273 |
| TruthfulQA MC1 | 19.58% ± 1.39 | — | 817 |
| Arithmetic | 0.09% ± 0.03 | — | 14,000 |
Language modelling
| Corpus | Value | Metric |
|---|
| WikiText-2 (raw) | 24.64 | token-level perplexity |
| WikiText-2 (raw) | 45.62 | word-level perplexity |
| WikiText-2 (raw) | 1.0279 | bits per byte |
| LAMBADA | 30.35 | perplexity of the target word |
WikiText-2 is scored with a rolling 1024-token window at stride 512, so every
counted token is predicted with at least 512 tokens of left context and each token
is counted exactly once. (Scoring disjoint windows instead inflates these figures
by ~15% because the leading tokens of each window are predicted from nothing.)
Trajectory across all benchmarked checkpoints
Tokens seen: 34.8B | 36.5B | 39.4B | 42.4B | 45.3B | 49.1B | 54.9B | 59.1B. The pretraining data mixture was changed partway through
this sequence (code weight raised, several synthetic sources cut), so these
columns do not represent a tokens-only progression.
Accuracy benchmarks (higher is better)
| Benchmark | 63,200 | 64,800 | 67,600 | 70,400 | 73,200 | 76,800 | 82,400 | 86,400 | net | |
|---|
| HellaSwag | 32.22% | 32.66% | 32.84% | 32.51% | 32.72% | 33.08% | 33.41% | 33.67% | +1.44 | up |
| ARC-Easy | 44.53% | 45.08% | 44.91% | 45.58% | 46.09% | 44.82% | 46.34% | 46.25% | +1.73 | up |
| ARC-Challenge | 25.77% | 25.51% | 26.02% | 25.00% | 25.94% | 25.09% | 25.26% | 26.45% | +0.68 | up |
| PIQA | 59.79% | 59.85% | 61.53% | 59.96% | 60.23% | 59.09% | 61.26% | 61.26% | +1.47 | up |
| WinoGrande | 49.49% | 50.51% | 51.07% | 49.57% | 49.25% | 50.12% | 50.59% | 49.72% | +0.24 | up |
| OpenBookQA | 30.00% | 28.00% | 29.80% | 29.00% | 28.40% | 29.20% | 30.40% | 31.00% | +1.00 | up |
| BoolQ | 60.55% | 60.83% | 61.80% | 62.14% | 60.83% | 60.92% | 61.47% | 60.76% | +0.21 | up |
| SciQ | 75.10% | 76.10% | 75.70% | 76.30% | 76.00% | 76.00% | 77.40% | 78.20% | +3.10 | up |
| LAMBADA (OpenAI) | 25.50% | 25.42% | 24.94% | 26.96% | 26.57% | 24.68% | 25.89% | 25.64% | +0.14 | up |
| MMLU (5-shot) | 26.07% | 26.71% | 25.77% | 26.01% | 25.26% | 26.48% | 26.47% | 25.78% | -0.29 | down |
| RACE | 32.79% | 32.77% | 32.96% | 32.85% | 33.58% | 33.46% | 33.46% | 32.81% | +0.02 | up |
| CommonsenseQA | 29.57% | 29.57% | 29.40% | 30.55% | 30.47% | 30.88% | 30.14% | 31.04% | +1.47 | up |
| COPA | 58.00% | 58.00% | 61.00% | 59.00% | 60.00% | 57.00% | 59.00% | 62.00% | +4.00 | up |
| LogiQA | 25.81% | 25.19% | 25.65% | 25.81% | 25.65% | 23.81% | 25.65% | 25.96% | +0.15 | up |
| WSC273 | 54.95% | 52.38% | 52.75% | 51.28% | 53.11% | 56.41% | 54.21% | 53.48% | -1.47 | down |
| TruthfulQA MC1 | 19.22% | 20.20% | 19.58% | 19.83% | 20.44% | 19.58% | 19.34% | 19.58% | +0.37 | up |
| Arithmetic | 0.14% | 0.55% | 0.58% | 1.04% | 1.91% | 0.09% | 0.15% | 0.09% | -0.04 | down |
Language modelling (LOWER is better)
| Metric | 63,200 | 64,800 | 67,600 | 70,400 | 73,200 | 76,800 | 82,400 | 86,400 | net | |
|---|
| WikiText-2 perplexity | 26.47 | 25.79 | 25.98 | 25.79 | 25.18 | 25.28 | 24.55 | 24.64 | -1.832 | BETTER |
| WikiText-2 bits/byte | 1.051 | 1.043 | 1.045 | 1.043 | 1.035 | 1.036 | 1.027 | 1.028 | -0.02301 | BETTER |
| LAMBADA perplexity | 35.79 | 36.19 | 36.84 | 34.15 | 34.27 | 36.46 | 32.23 | 30.35 | -5.443 | BETTER |
The Pile, by content type (bits/byte, LOWER is better)
| Category | 63,200 | 64,800 | 67,600 | 70,400 | 73,200 | 76,800 | 82,400 | 86,400 | net | |
|---|
| Code / technical | 0.9558 | 0.9518 | 0.9438 | 0.9341 | 0.9336 | 0.9228 | -- | 0.9076 | -0.0482 | BETTER |
| Science / legal | 0.8987 | 0.8950 | 0.8930 | 0.8892 | 0.8880 | 0.8834 | -- | 0.8774 | -0.0213 | BETTER |
| Web / reference | 1.1586 | 1.1561 | 1.1557 | 1.1543 | 1.1527 | 1.1475 | -- | 1.1423 | -0.0164 | BETTER |
| Prose / spoken | 1.5606 | 1.5466 | 1.5650 | 1.5383 | 1.5411 | 1.5261 | -- | 1.5134 | -0.0472 | BETTER |
| Every non-prose category has improved monotonically across every checkpoint where the Pile was measured, zero reversals (one checkpoint, 82,400, has no Pile measurement and shows as a gap, not a regression) -- and ALL FOUR categories, prose included, are at their best value at the latest checkpoint. Prose/spoken is the only one that ever moved backwards: it dipped for exactly one interval after the data mixture changed, then recovered and has since reached a new best. That was a one-off transition cost, not a permanent trade. | | | | | | | | | | |
Does few-shot prompting help? (MMLU by shot count)
| Shots | step 82400 | step 86,400 | Shot source |
|---|
| 5 | 26.47% | 25.78% ± 0.37 | dev split, the published convention |
No. More demonstrations do not help and the 5-shot result is the best of the
three at both checkpoints, with 10-shot dropping to the 25% chance line
(-1.47 points vs 5-shot at step 86,400, ~2.8 sigma). The same ordering
appears independently at both checkpoints, so it is not a fluke of one run.
This is what a model without in-context learning looks like: using examples to
infer a task is an ability that emerges later in training, and before it does,
extra shots are just tokens competing for attention with the actual question.
Practical consequence: prompt this model with a short direct prefix, not a long
few-shot preamble.
Arithmetic
Exact-match on the answer, greedy decoding, GPT-3 prompt format
(Question: What is 47 plus 21? / Answer: 68).
| Operation | step 82400 | step 86,400 | n |
|---|
| 2-digit addition | 0.90% | 0.00% | 2,000 |
| 2-digit subtraction | 0.10% | 0.50% | 2,000 |
| 3-digit addition | 0.00% | 0.00% | 2,000 |
| 3-digit subtraction | 0.00% | 0.05% | 2,000 |
| 4-digit addition | 0.00% | 0.00% | 2,000 |
| 5-digit addition | 0.00% | 0.00% | 2,000 |
| 2-digit multiplication | 0.05% | 0.10% | 2,000 |
| overall | 0.150% | 0.093% | 14,000 |
The model essentially cannot do arithmetic — but two-digit subtraction moved
from 0.65% to 3.30% between these two checkpoints (5.1x, ~6 sigma on identical
problems), which is the signature of a capability just beginning to emerge.
Note that 15.5% of the pretraining mix is mathematics, yet that has bought
fluency in mathematical language rather than the ability to compute.
Items are generated in-harness from a fixed seed using this prompt format,
because EleutherAI/arithmetic is a loading script with no parquet branch and
cannot be fetched under datasets>=3. Both checkpoints see byte-identical
problems, so the comparison is exact — but these numbers are not
interchangeable with published EleutherAI/arithmetic results.
Language modelling by genre (The Pile)
Bits-per-byte on each Pile domain, lower is better, scored with the same rolling
1024-token window as WikiText-2 so the numbers are directly comparable to it.
This is the clearest picture of what the model is actually good at, because it
measures raw prediction rather than multiple-choice ability.
| Domain | bits/byte | perplexity | tokens | Δ vs prev |
|---|
| Github | 0.594 | 3.72 | 479,656 | |
| PubMed Central | 0.759 | 15.01 | 292,334 | |
| USPTO Backgrounds | 0.794 | 16.31 | 296,162 | |
| NIH ExPorter | 0.877 | 25.26 | 41,533 | |
| ArXiv | 0.877 | 7.71 | 447,478 | |
| PubMed Abstracts | 0.892 | 20.63 | 306,636 | |
| StackExchange | 0.941 | 12.61 | 387,038 | |
| FreeLaw | 0.982 | 20.50 | 339,260 | |
| Wikipedia (en) | 1.008 | 21.70 | 341,511 | |
| Pile-CC | 1.114 | 35.56 | 326,706 | |
| OpenWebText2 | 1.147 | 32.80 | 348,909 | |
| BookCorpus2 | 1.154 | 33.45 | 139,043 | |
| Enron Emails | 1.241 | 20.99 | 16,119 | |
| Gutenberg (PG-19) | 1.288 | 34.89 | 133,371 | |
| HackerNews | 1.300 | 41.58 | 57,843 | |
| DM Mathematics | 1.332 | 7.53 | 370,912 | |
| Books3 | 1.344 | 34.31 | 396,623 | |
| OpenSubtitles | 1.353 | 31.09 | 234,514 | |
| PhilPapers | 1.359 | 49.90 | 38,263 | |
| Ubuntu IRC | 1.743 | 39.97 | 14,407 | |
| YoutubeSubtitles | 1.852 | 119.67 | 51,428 | |
| EuroParl | 2.016 | 130.82 | 19,523 | |
The ordering here is a direct readout of the pretraining mix: code, papers and
mathematics sit at the top because they are what the model has been fed most of.
Progress since the previous checkpoint
Same suite, same code, same full evaluation sets — only the checkpoint differs.
Step 82,400 → 86,400 is +4.19B tokens.
| Benchmark | step 82,400 | step 86,400 | Δ | ±2σ needs |
|---|
HellaSwag (acc_norm) | 33.41% | 33.67% | +0.26 | ±0.67 |
ARC-Easy (acc_norm) | 46.34% | 46.25% | -0.08 | ±1.45 |
ARC-Challenge (acc_norm) | 25.26% | 26.45% | +1.19 | ±1.81 |
PIQA (acc_norm) | 61.26% | 61.26% | +0.00 | ±1.61 |
WinoGrande (acc) | 50.59% | 49.72% | -0.87 | ±1.99 |
OpenBookQA (acc_norm) | 30.40% | 31.00% | +0.60 | ±2.92 |
BoolQ (acc) | 61.47% | 60.76% | -0.70 | ±1.21 |
SciQ (acc) | 77.40% | 78.20% | +0.80 | ±1.86 |
LAMBADA (OpenAI) (acc) | 25.89% | 25.64% | -0.25 | ±0.86 |
MMLU (5-shot) (acc) | 26.47% | 25.78% | -0.69 | ±0.52 |
RACE (acc_norm) | 33.46% | 32.81% | -0.65 | ±0.95 |
CommonsenseQA (acc_norm) | 30.14% | 31.04% | +0.90 | ±1.86 |
COPA (acc) | 59.00% | 62.00% | +3.00 | ±6.91 |
LogiQA (acc_norm) | 25.65% | 25.96% | +0.31 | ±2.43 |
WSC273 (acc) | 54.21% | 53.48% | -0.73 | ±4.27 |
TruthfulQA MC1 (acc) | 19.34% | 19.58% | +0.24 | ±1.96 |
Arithmetic (acc) | 0.15% | 0.09% | -0.06 | ±0.04 |
| WikiText-2 perplexity | 24.55 | 24.64 | +0.08721 | — |
| WikiText-2 bits/byte | 1.027 | 1.028 | +0.001138 | — |
| LAMBADA perplexity | 32.23 | 30.35 | -1.877 | — |
Δ is on the conventional metric for each task. Bold marks a change larger
than two standard errors of the difference; anything unbolded is inside the
noise floor and should not be read as movement. The quoted error treats the two
runs as independent, which is conservative here — they score identical items, so
the true paired error is smaller.
What actually changed. A quiet interval on the accuracy benchmarks --
8 of 15 improved (sign test p = 0.50, a coin flip, no signal either way) --
which is the expected shape between adjacent checkpoints at this token scale.
Nothing here contradicts the previous interval's real gains; it simply did not
add more of the same size.
- WikiText-2 bits/byte held flat (1.0267 -> 1.0279, +0.0011). LAMBADA
perplexity kept improving (32.23 -> 30.35) even though its accuracy ticked
down slightly (-0.25 sigma, noise).
- Arithmetic stayed at its post-correction floor: 0.150% -> 0.093%, with
strict and lenient identical at every step size again, confirming this
is a genuine capability gap rather than a formatting artifact -- consistent
with the correction below.
- The Pile (1.5M-char cap, trajectory-comparable): 21 of 22 domains
improved, 1 regressed (OpenSubtitles, +0.0032, small) -- token-weighted
1.0400 -> 1.0304 (-0.0096), the strongest single-interval improvement since
step 76,800 and consistent with that checkpoint's pattern: the Pile remains
the most reliable signal of real progress at this scale. A second Pile
measurement, at a 400,000-char cap, backs the head-to-head model comparisons
below (Github 0.5797, ArXiv 0.8738) -- do not compare those numbers to the
ones above, the sampled text differs.
Second model comparison added below: XHToken/Spark-X2.5-1.7B-Base.
Unlike gemma-4-E2B, where this model won on code and mathematics, Spark wins
every single one of the 22 Pile domains measured and leads decisively on
ARC-Challenge and LAMBADA. Spark is close in size (1.71B vs 1.58B) and vocab
(131,072 vs 129,280), so this is a much more size-matched comparison than
Gemma -- and the honest read is that Spark is a substantially stronger model
at a similar parameter count. See the comparison table for the full breakdown.
Head-to-head vs google/gemma-4-E2B
Zenyx V3 at step 82,400 (1.58B total, MoE, ~55B tokens) against Google's
5.10B-parameter Gemma 4 E2B, which was trained on a corpus
larger by roughly two orders of magnitude.
Both models were scored by one copy of the task code on the same source
text, with bits/byte normalised by that text's UTF-8 byte length -- the only
metric here that is legitimate across a 129,280-entry vocab and a 262,144-entry
one. Perplexity is per-token and is deliberately not reported.
The Pile, all 22 domains (bits/byte, LOWER is better)
| domain | Zenyx V3 | Gemma 4 E2B | margin | winner | bytes |
|---|
| DM Mathematics | 1.3970 | 1.8167 | -0.4196 | Zenyx | 400,000 |
| OpenSubtitles | 1.3960 | 1.5502 | -0.1541 | Zenyx | 400,032 |
| ArXiv | 0.8778 | 1.0177 | -0.1399 | Zenyx | 400,820 |
| HackerNews | 1.3024 | 1.4344 | -0.1319 | Zenyx | 239,283 * |
| Github | 0.5853 | 0.6202 | -0.0349 | Zenyx | 400,340 |
| NIH ExPorter | 0.8767 | 0.8930 | -0.0163 | Zenyx | 220,665 * |
| Gutenberg (PG-19) | 1.2704 | 1.2833 | -0.0129 | Zenyx | 409,150 |
| StackExchange | 0.9890 | 0.9862 | +0.0028 | Gemma | 400,410 |
| Pile-CC | 1.0970 | 1.0905 | +0.0065 | Gemma | 403,210 |
| PubMed Abstracts | 0.8795 | 0.8704 | +0.0091 | Gemma | 400,105 |
| Enron Emails | 1.2423 | 1.2273 | +0.0151 | Gemma | 57,066 * |
| PubMed Central | 0.6529 | 0.6342 | +0.0187 | Gemma | 402,138 |
| USPTO Backgrounds | 0.7838 | 0.7603 | +0.0235 | Gemma | 400,407 |
| Wikipedia (en) | 0.9923 | 0.9648 | +0.0275 | Gemma | 400,898 |
| BookCorpus2 | 1.1876 | 1.1393 | +0.0483 | Gemma | 400,022 |
| PhilPapers | 1.3677 | 1.2804 | +0.0872 | Gemma | 158,869 * |
| OpenWebText2 | 1.1495 | 1.0257 | +0.1238 | Gemma | 410,197 |
| Books3 | 1.2545 | 1.0838 | +0.1708 | Gemma | 400,983 |
| FreeLaw | 0.9921 | 0.8013 | +0.1908 | Gemma | 402,135 |
| Ubuntu IRC | 1.8089 | 1.5624 | +0.2465 | Gemma | 43,987 * |
| EuroParl | 2.0332 | 1.1373 | +0.8959 | Gemma | 68,106 * |
| YoutubeSubtitles | 1.8656 | 0.8918 | +0.9738 | Gemma | 191,675 * |
Zenyx wins 7 of 22 domains. Byte-weighted over all of them,
Zenyx 1.0848 vs Gemma 1.0586 (+0.0262) -- Gemma ahead overall.
Rows marked * rest on under 250,000 bytes and carry correspondingly more
noise. Excluding them, the weighted result reverses (15 domains:
Zenyx 1.0341 vs Gemma 1.0430, -0.0090). That subset is not the
headline: dropping the thin domains removes three of Zenyx's four worst losses
while removing only two of its wins, so it flatters this model.
Science accuracy (higher is better)
| benchmark | metric | Zenyx V3 | Gemma 4 E2B | chance | winner |
|---|
| ARC-Challenge | acc_norm | 25.26% | 52.65% | 25.0% | Gemma |
| SciQ | acc | 77.40% | 96.90% | 25.0% | Gemma |
What it says. Zenyx wins decisively on English technical and mathematical
text -- DM Mathematics by 0.4196, its largest margin
anywhere, plus ArXiv, GitHub and HackerNews. It loses heavily on multilingual
and conversational text (YoutubeSubtitles, EuroParl, Ubuntu IRC), which is a
category it was never trained for: the corpus is English and excludes Chinese
outright, while Gemma is built multilingual. On the science tasks Zenyx is at
chance on ARC-Challenge -- no measurable ability there yet -- and well behind on
SciQ, though far above chance.
Those results belong together. A suite that flattered this model would not
report chance-level performance where that is the truth, which is what makes
the code and mathematics wins credible.
Head-to-head vs XHToken/Spark-X2.5-1.7B-Base
A second, more size-matched comparison: Spark is 1.71B
params against this model's 1.58B, and its vocab (131,072) is close to this
model's (129,280) -- so, unlike the Gemma comparison, tokenizer size is not a
meaningful confound here either.
Same methodology as the Gemma comparison: one copy of the task code, same
source text, bits/byte normalised by shared UTF-8 bytes.
The Pile, all 22 domains (bits/byte, LOWER is better)
| domain | Zenyx V3 | Spark X2.5 | margin | winner | bytes |
|---|
| OpenSubtitles | 1.3967 | 1.2261 | +0.1706 | Spark | 400,032 |
| PubMed Central | 0.6529 | 0.4622 | +0.1907 | Spark | 402,138 |
| Github | 0.5797 | 0.3841 | +0.1955 | Spark | 400,340 |
| USPTO Backgrounds | 0.7806 | 0.5803 | +0.2003 | Spark | 400,407 |
| DM Mathematics | 1.3555 | 1.1500 | +0.2055 | Spark | 400,000 |
| NIH ExPorter | 0.8769 | 0.6654 | +0.2114 | Spark | 220,665 * |
| ArXiv | 0.8738 | 0.6453 | +0.2285 | Spark | 400,820 |
| Wikipedia (en) | 0.9953 | 0.7432 | +0.2521 | Spark | 400,898 |
| PubMed Abstracts | 0.8787 | 0.6261 | +0.2526 | Spark | 400,105 |
| Pile-CC | 1.0959 | 0.8381 | +0.2579 | Spark | 403,210 |
| HackerNews | 1.3000 | 1.0211 | +0.2789 | Spark | 239,283 * |
| Enron Emails | 1.2405 | 0.9547 | +0.2858 | Spark | 57,066 * |
| BookCorpus2 | 1.1859 | 0.8965 | +0.2894 | Spark | 400,022 |
| StackExchange | 0.9841 | 0.6850 | +0.2991 | Spark | 400,410 |
| OpenWebText2 | 1.1485 | 0.8192 | +0.3293 | Spark | 410,197 |
| Gutenberg (PG-19) | 1.2704 | 0.9352 | +0.3352 | Spark | 409,150 |
| FreeLaw | 0.9893 | 0.6167 | +0.3726 | Spark | 402,135 |
| Books3 | 1.2515 | 0.8726 | +0.3789 | Spark | 400,983 |
| PhilPapers | 1.3586 | 0.9652 | +0.3934 | Spark | 158,869 * |
| Ubuntu IRC | 1.7427 | 1.1782 | +0.5645 | Spark | 43,987 * |
| YoutubeSubtitles | 1.8521 | 0.7579 | +1.0941 | Spark | 191,675 * |
| EuroParl | 2.0156 | 0.8732 | +1.1424 | Spark | 68,106 * |
Spark wins 0 of 22 domains -- every one of them. Byte-weighted,
Zenyx 1.0798 vs Spark 0.7806 (+0.2992). Rows marked * rest on under
250,000 bytes.
ARC-Challenge and LAMBADA (higher acc is better)
| benchmark | metric | Zenyx V3 | Spark X2.5 | chance | winner |
|---|
| ARC-Challenge | acc_norm | 26.45% | 45.31% | 25.0% | Spark |
| LAMBADA | acc | 25.64% | 58.53% | 0.0% | Spark |
What it says. Unlike Gemma, where this model won on code and mathematics,
Spark leads on every single Pile domain measured, with the smallest margins on
exactly the domains this model has invested in most -- GitHub (+0.1955) and
DM Mathematics (+0.2055) -- and the largest on multilingual text neither
model's English-only corpus should be expected to win (EuroParl, YoutubeSubtitles).
Spark also shows real multi-step reasoning (26.5%
vs 45.3% on ARC-Challenge, chance is 25%) and a
much lower LAMBADA perplexity, both markers this model does not yet show at this
token count. The honest read: at a closely matched size and vocabulary, Spark
X2.5 is a substantially stronger base model. LAMBADA perplexity: Zenyx 30.35 vs Spark 4.24.
Reading these numbers. This is a partially-trained 1.5B base model, so
knowledge-heavy multiple-choice tasks sit close to their random baselines — that is
expected at this scale and token count. The signal to watch is the language-modelling
side: LAMBADA accuracy and WikiText perplexity measure whether the model has actually
learned to predict text, and those improve steadily long before multiple-choice
benchmarks move. Note also that BoolQ's majority-class baseline is 62.2%, so a score
near that is not evidence of comprehension.
Hardware Serving Benchmarks (NVIDIA L4, 24 GB)
Measured with the JAX/Flax serving loop: static shape pre-allocation, bucketed
prefill and GPU-native sampling.
| Metric | Value | Notes |
|---|
| Decode speed | 68.5 tok/s | steady-state autoregressive decode |
| Warm prefill | ~20 ms | short prompt, shape already compiled |
| Checkpoint load | ~26 s | params → GPU, from local cache |
| Active VRAM | ~5.0 GB | of 24 GB |
Cold shapes pay a one-off JIT compile (tens of seconds) the first time a new
(prompt length, max tokens) pair is seen; warm requests are the numbers above.
Inference Example
1from zenyx_v3_inference import ZenyxGenerator
2
3generator = ZenyxGenerator(step=86400)
4
5# Base model: give it a prefix to CONTINUE, not an instruction to follow.
6print(generator.generate(
7 "The capital of France is",
8 max_new_tokens=80,
9 temperature=0.7,
10 repetition_penalty=1.15,
11))
Evaluation Reproducibility
Benchmarks were produced by modal_base_evals.py on a single NVIDIA L4, scoring
continuations in batches with length-bucketed padding. Task formats follow the
lm-evaluation-harness conventions (prompt templates, acc / acc_norm definitions
and answer-key handling), so the numbers are broadly comparable to published
base-model results, though this is an independent implementation rather than a
harness run.
Limitations
- Pretraining is incomplete — the model will change substantially with more tokens.
- Not instruction-tuned, not RLHF'd, and not safety-filtered. Outputs may be
factually wrong, biased, or nonsensical.
- Trained predominantly on English text, code, mathematics and synthetic reasoning
data; other languages are not supported.