A 94.8M-parameter decoder (12 layers, d=768, GQA 12/4, SwiGLU, RoPE, tied embeddings)
trained from scratch on 2.75B tokens of a corpus built and filtered from source:
verified mathematics, reasoning traces, proofs, code and Italian.
ONNX, dynamically quantised to uint8 (~120 MB), so it runs in a browser through
transformers.js with no server.
Chat format:
<|user|>
{{message}}
<|assistant|>
Trained on 2x RTX 5090 at 302k tokens/s (57% MFU).