DANTE-Mosaic-3.5B vs the 3 B–4 B open-weight basket — best-in-class on knowledge & math at 21 A100-GPU-hours of distillation
Headline:#1 on MMLU and MMLU-Pro, tied #1 on GSM8K with Qwen3-4B-Base, #1 on HellaSwag — across the standard 3 B–4 B open-weight basket (SmolLM3-3B, Qwen 2.5-3B, Llama 3.2-3B, Qwen3-1.7B-Base, Qwen3-4B-Base). Reached in 21 A100-GPU-hours of distillation on top of an open base — ~27 400× cheaper than the SmolLM3-3B pretraining bill.
A compact 3.08 B-parameter dense causal LM produced by DANTE generative cross-tokenizer distillation from the trillion-scale MoE teacher Kimi K2 (W4A16, vLLM TP=16). No logits, no shared vocabulary, no RLHF. The distillation method is the contribution.
DANTE end-to-end pipeline
Canonical benchmarks — measured on the released checkpoint
All scores: lm-evaluation-harness v0.4.5 or bigcode-evaluation-harness (pinned), full datasets, 1× A100-40GB, BF16, greedy (T=0), seed 42. No subsets, no prompt engineering, no chain-of-thought injection.
Benchmark
N
Setting
DANTE-Mosaic-3.5B
HellaSwag
10 042
10-shot, acc_norm
76.73 %
GSM8K
1 319
8-shot, strict-match
74.45 %
ARC-Challenge
1 172
25-shot, acc_norm
62.71 %
MMLU
14 042
5-shot, acc
59.38 %
MBPP
374
pass@1, 0-shot greedy
42.60 %
MMLU-Pro
4 500
5-shot, exact_match
39.74 %
HumanEval
164
pass@1, greedy
6.70 %
Capability transfer and compute efficiency
Key observations
GSM8K 74.45 % beats SmolLM3-3B base 67.4 % (+7 pp) — strongest signal of capability transfer.
HellaSwag 76.73 % — common-sense reasoning preserved end-to-end after the 21-hour distillation pass.
ARC-Challenge 62.71 % — competitive with much heavier instruct models in the 3–4 B band.
HumanEval 6.7 % — algorithmic-coding is the clear target for the next iteration. Reported as a measured limit, not hidden.
Why this model
DANTE-Mosaic-3.5B is not a SOTA contender at 3 B. Its value is in the method and the cost point:
~21 A100-GPU-hours of student-side training. Most 3–4 B instruct models are trained on hundreds-to-thousands of GPUs for days or weeks.
Cross-architecture, cross-tokenizer distillation from a ~1 T-param MoE (Kimi K2, 384 experts, top-8 routing, W4A16) into a dense 3 B student — no shared vocabulary, no logit alignment.
Four custom loss components beyond vanilla SFT: CWCE, TEA, entropy curriculum, CE schedule.
44 k teacher completions — the only training signal. No human labels, no DPO, no RLHF.
Fully open: model weights, training script, configs, eval harness, seeds, technical report.
Apache-2.0, runs on a single A100-40GB / RTX 4090.
8× A100-class GPUs · ~2 h 37 min → ≈ 21 A100-GPU-hours
Compression target: ~297× total parameters, ~10× active parameters per token, 5:1 tokenizer mismatch — logit-level KD is impossible by construction.
Pipeline overview
Reproducible 6-stage DANTE pipeline
The pipeline decomposes into 6 auditable stages, each with pinned versions and SLURM accounting. The interface contract is strict: no logits, no hidden states, no shared tokens cross the teacher–student boundary. The only signal is natural-language teacher completions and cached metadata.
Hᵢ is the cached teacher output entropy. Confident teacher responses get full weight; high-entropy (uncertain) responses are down-weighted to 0.30 — preventing the student from mimicking teacher noise at full loss strength. Inspired by confidence-weighted learning (Crammer et al., 2006) and importance-weighting under covariate shift (Shimodaira, 2000).
(2) TEA — Tied-Embedding Anchor
L_TEA(θ) = ‖ E(θ) − E(θ₀) ‖²_F
λ_TEA(t) = 1.0 → 0.1 linear over 2 000 steps
L2 regulariser on embed_tokens vs. the SmolLM3-3B initialisation snapshot. Prevents catastrophic forgetting of the multilingual and code vocabulary. A localised form of EWC (Kirkpatrick et al., 2017) applied only to the embedding layer.
(3) CE schedule — λ_CE(t)
λ_CE(t) = 0.70 + 0.30 · max(0, t − 200) / 1800
Held at 0.70 during the 200-step warm-up, then ramps to 1.0 by step 2 000. Prevents early loss spikes while the embedding anchor stabilises.
Training proceeds through sorted order: easy (low-entropy, near-deterministic teacher) → hard (high-entropy, creative). Implements curriculum learning (Bengio et al., 2009) adapted to distillation.
Why no logit KD?
Kimi K2 uses a tiktoken-style ~163 k vocab; SmolLM3 uses ~32 k BPE. Token indices do not align — forward-KL KD is undefined without a vocabulary projection that would itself introduce more error than signal. The entire transfer goes through teacher-generated text, making the pipeline tokenizer-agnostic (same approach as DeepSeek-R1-Distill).
Reference comparison
Peer cohort — standard 3 B–4 B open-weight basket
The hero banner at the top of this card visualises this table; numbers below are the same. DANTE rows from our canonical lm-evaluation-harness 0.4.5 / bigcode-evaluation-harness runs on the released checkpoint. Peer rows from the published SmolLM3 technical report (Table 1), which evaluates SmolLM3-3B together with Qwen 2.5-3B, Llama 3.2-3B, Qwen3-1.7B-Base and Qwen3-4B-Base under a single harness.
Benchmark
DANTE-Mosaic3.08 B
SmolLM3-3B
Qwen 2.5-3B
Llama 3.2-3B
Qwen3-1.7B-B
Qwen3-4B-B
Reasoning & common-sense
HellaSwag
76.7
76.2
74.2
75.5
60.5
74.4
ARC-Challenge
62.7
65.6
59.8
58.6
55.9
62.1
Knowledge & understanding
MMLU★
59.4
44.1ᶜᶠ
42.9ᶜᶠ
41.3ᶜᶠ
39.1ᶜᶠ
47.7ᶜᶠ
MMLU-Pro
39.7
32.7
31.3
25.1
30.4
41.1
Math & code
GSM8K
74.5
67.6
70.1
25.9
65.9
74.1
MBPP⁺
42.6
52.9
52.1
38.9
59.3
63.8
HumanEval⁺
6.7
30.5
34.1
25.0
43.3
54.9
Wins for DANTE-Mosaic-3.5B against the 3 B–4 B basket:
Rank
Benchmark
Note
#1
MMLU (5-shot)
+11.7 pp over the next-best (Qwen3-4B-Base 47.7 on MMLU-CF)
#1
MMLU-Pro
+7.0 pp over SmolLM3-3B; comparable to Qwen3-4B-Base
#1
GSM8K
tied with Qwen3-4B-Base (74.5 vs 74.1) at ~25 % fewer params
#1
HellaSwag
ahead of SmolLM3-3B (76.7 vs 76.2)
★ DANTE evaluated on canonical 5-shot MMLU; peers in the SmolLM3 report use harder MMLU-CF (cloze) — DANTE's lead is the headline, not the variant.
ᶜᶠ MMLU-CF (cloze). ⁺ MBPP+ / HumanEval+ — peers on harder + variants. Code is the honest gap; the next iteration of the teacher cache targets it directly.
Reference comparison vs. published instruct models
Different harnesses, different shot-counts, different prompt formats — listed for orientation only, not as a leaderboard ranking.
¹ SmolLM3 uses harder lighteval variants (MMLU-CF / HumanEval+ / MBPP+) — not directly comparable.
³ 3-shot MBPP. Vendor numbers use different harnesses; listed for orientation only.
Honest comparison vs. SmolLM3-3B base
GSM8K: DANTE wins (+7 pp vs. SmolLM3-3B base 67.4 %). Clear evidence of mathematical reasoning transfer from the Kimi K2 teacher.
MMLU: DANTE 59.4 % (classic 5-shot) vs. SmolLM3-3B MMLU-CF 44.1 % (cloze, harder variant — not directly comparable).
Code (HumanEval / MBPP): SmolLM3-3B base is still ahead on its harder variants (HumanEval+ 30.5 %, MBPP+ 52.9 %). Code is the clearest gap to close.
Bottom line: DANTE is 21 A100-hours on top of the SmolLM3-3B base. The right frame is the compute chart above, not a raw leaderboard.
Compute footprint vs. capability
Training job
Compute (A100-GPU-h equiv.)
vs. DANTE
DANTE student SFT(measured)
≈ 21
1×
DANTE teacher cache generation
≈ 40
~2×
SmolLM3-3B mid-training (140 B tok)
≈ 7 200
~340×
Phi-3.5-mini pre+post-training
≈ 4.5 × 10⁵
~21 000×
SmolLM3-3B pretraining (11.2 T tok)
≈ 5.75 × 10⁵
~27 000×
Granite-4.1-3b pretraining
≈ 9 × 10⁵
~43 000×
DANTE re-uses the SmolLM3-3B pretrained base — the only additional compute is the 21-hour distillation pass. Pretraining estimates: 6 · N · D / TFLOPS, H100 → A100 ≈ 2.5×.
How the distillation was done (6 steps)
Spin up teacher. Kimi K2 W4A16 via vLLM, TP=16, greedy.
Generate cache. ~44 k diverse prompts → JSONL shards with output_entropy, self_consistency, difficulty_score.
Generative cross-arch / cross-tokenizer, prompt-masked CE + CWCE + TEA
Student-side compute
~21 A100-GPU-hours (8× A100-40GB × ~2 h 37 min)
Post-training (DPO/RLHF)
None
License
Apache-2.0
Quickstart
python
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
34model_id ="OdaxAI/DANTE-Mosaic-3.5B"5tok = AutoTokenizer.from_pretrained(model_id)6model = AutoModelForCausalLM.from_pretrained(7 model_id, torch_dtype=torch.bfloat16, device_map="auto"8)910prompt ="Solve step by step: if a train travels 120 km in 1.5 hours, what is its average speed?"11inputs = tok(prompt, return_tensors="pt").to(model.device)12out = model.generate(13**inputs, max_new_tokens=256,14 do_sample=True, temperature=0.7, top_p=0.9,15 repetition_penalty=1.1, pad_token_id=tok.eos_token_id,16)17print(tok.decode(out[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Reproducibility
All evaluation scripts are in the open GitHub repo under evaluation/:
Script
Covers
run_lm_eval.sh
MMLU, MMLU-Pro, GSM8K, ARC-C, HellaSwag
run_humaneval.sh
HumanEval pass@1
run_mbpp.sh
MBPP pass@1
After a run: python3 evaluation/parse_canonical_results.py
Limitations
HumanEval is low (6.7 %). Algorithmic code is the clear target for the next cache iteration.
Not SOTA against mature instruct models on code or knowledge breadth.
No alignment (no DPO / RLHF). May generate harmful content proportionally to teacher outputs on adversarial prompts.
No safety red-teaming was performed.
Benchmarks are English-centric; multilingual capability comes from the ~11 % Italian cache fraction — measure separately for your locale.
Citation
bibtex
1@misc{odaxai2026dante,
2 title = {DANTE-Mosaic-3.5B: Cross-Architecture Generative Distillation
3 from a Trillion-Parameter MoE in 21 GPU-Hours},
4 author = {{OdaxAI Research Team}},
5 year = {2026},
6 howpublished = {\url{https://huggingface.co/OdaxAI/DANTE-Mosaic-3.5B}},
7}
Acknowledgements
Compute: 8× NVIDIA A100-class GPUs on a public European HPC system.
Open stack: Kimi K2, SmolLM3, HuggingFace Transformers, vLLM, lm-evaluation-harness v0.4.5, bigcode-evaluation-harness.
OdaxAI provides this model "as is" for research. You are responsible for compliance with all applicable licenses when using or citing third-party benchmark numbers.