TinyGPT-ML-84M (TGPT-XL)
An 85M parameter language model trained entirely from scratch — no pretrained base, no borrowed weights, no fine-tuning on top of someone else's model.
Architecture
Modern Llama-style stack: RoPE, RMSNorm, SwiGLU
ChatML format with a dedicated stop token
Trained on 8B tokens (FineWeb-Edu + SmolTalk), heavily overtrained past Chinchilla-optimal for maximum capability-per-parameter
Benchmarks
Evaluated head-to-head against Supra-1.5-50M-Instruct across 5 categories (42 questions, fair decode settings: temp 0.3, top-p 0.9, repetition penalty 1.15).
Category
TGPT-XL
Supra-50M
Knowledge
50%
40%
Retention
83%
50%
Reasoning
37%
12%
Advice
100%
38%
Overall
62%
36%
TGPT-XL wins or ties every single category, with its biggest edges in multi-turn retention (holds context across turns) and open-ended advice (structured, actionable responses). Supra tends to deflect on emotional/open-ended prompts.
Highlights
Clean multi-turn context recall
Respects length and format constraints
Reliable stop behavior
Runs on-device (Pixel 6a) via GGUF