Views
No views yet
| Parameters | 5,868,800 |
| Total bytes (packed, spec §6) | 11,730,944 (11.19 MiB) |
| Layers / width / heads | 6 / 256 / 8 |
| Context | 512 |
| Vocab | 4,096 (custom BPE, ByteLevel) |
| In-block linear weights | fp16 (2 bytes/weight) |
| Embeddings / LM head (tied) | fp16 |
| Seeds | 0, 1 (both published, see below) |
| Training tokens | 500M |
model.safetensors (fp16). Ternary in-block linears are stored
dequantized — scale * {-1,0,+1} already materialized in fp16 — so both arms load
into the same fp16 model class.1from safetensors.torch import load_file
2from nanofable.config import TIERS
3from nanofable.model import build_model
4from nanofable.tokenizer import load_tokenizer
5from nanofable.generate import generate
6
7model = build_model(TIERS["small"], "fp16")
8model.load_state_dict(load_file("model.safetensors"), strict=False) # lm_head re-ties to tok_emb
9model.eval()
10
11tok = load_tokenizer("tokenizer.json")
12print(generate(model, tok, "Once upon a time", temperature=1e-4, top_k=0))model.safetensors # seed 0 — the weights this card describes
config.json
model.tpack # packed browser payload (what the demo site streams)
meta.json eval.json metrics.csv
tokenizer.json # 4,096-vocab custom BPE
seed1/ # second training replica, same files| seed 0 (root) | seed 1 | |
|---|---|---|
| Val PPL | 8.83 | 8.70 |
| Judge (0-5, n=200) | 2.413 [2.29, 2.54] | 2.345 [2.24, 2.45] |
model.tpack is the packed browser payload (pack format v1), fp16 throughout. This is
the artifact the byte accounting above refers to, and the one the demo site streams for
in-browser inference. It carries the same weights as model.safetensors, in the container
the browser runtime reads.