Note: SFT was not done on this model, it cannot respond to questions 9/10 times.
Glint-1.3
⚠️ IMPORTANT NOTICE
This model is experimental. Glint-1.3 is a 982K parameter research model.
Performance characteristics: This model may occasionally output chuamliamce. If it does, try again. It is shy.
Not production-ready: This is a tiny neural network running on a prayer and a GPU.
Quick Stats
Stat
Value
Parameters
982,656 (under 1M 👍)
Training Tokens
100 Billion (FineWeb-Edu)
Hardware
RTX 5090
Context Window
256 tokens
Inference Speed
138,562.1925 tok/s
Vibe
Doing its best
What Is This?
Glint-1.3 is the first model in the CompactAI scaling-down plan.
We spent months adding features, SPIN, DPO, sleep gates, retention, recurrent loops, LoRA, engrams. More parameters, more tricks, more complexity. And you know what? The features were hurting the models. The tiny models couldn't breathe. So we're doing the opposite now: scaling down. Strip everything. Pure Llama. See how far simplicity goes.
This is that experiment. ~1M params. No gimmicks. Just a transformer doing its best.
It runs at 138,000 tokens per second on an RTX 5090. Fun, but useless. lmao.
The Journey
The model improves monotonically over 95K training steps on 100B tokens, with Wikitext-2 cross-entropy loss dropping from 4.29 → 3.08. For a 1M parameter model, this is actually respectable.
Model Specifications
Parameter
Value
Architecture
Transformer Decoder (Llama-style)
Parameters
982,656
Hidden Dim
128
Layers
4
Attention Heads
4
KV Heads
4 (GQA)
MLP Intermediate
384 (SwiGLU)
Context Length
256 tokens
Vocab Size
500 (ByteLevel BPE)
Normalization
RMSNorm
Position Encoding
RoPE
Embeddings
Tied input/output
Benchmarks
All checkpoints evaluated on Wikitext-2, BLiMP (grammaticality), and ARC-Easy (science QA). Sliding-window log-prob scoring methodology from the CompactAI benchmark suite.
Checkpoint Benchmark
Per-Metric Standouts
Metric
Best Checkpoint
Score
Wikitext-2 CE Loss
Step 95,000
3.06
BLiMP Accuracy
Step 11,500
64.2%
ARC-Easy Accuracy
Step 55,500
32.5%
Merged Model (Model Soup)
Weight averaging the best checkpoints per benchmark via per-parameter-group SLERP produces a model that exceeds individual bests on certain metrics:
Model
WT Loss
BLiMP
ARC
Composite
Best Merged
3.148
68.7% 🏆
29.0%
1.391 🏆
Best WT (step 95367)
3.080
53.7%
25.0%
1.431
Best BLiMP (step 11500)
3.307
64.2%
22.5%
1.480
Best ARC (step 55500)
3.128
50.7%
32.5%
1.432
The merged model achieves superadditive BLiMP gains (+4.5% over the individual best checkpoint) through spherical interpolation of attention and MLP weights at different blend factors.