Views
No views yet
esm2 architecture) — 7.4M (7,422,080) parameters, trained on 89.1M (89,064,960) UniRef50 residues.nanoprot-esm2-XS is part of the nanoprot suite: a Pythia-style matrix of protein
language models spanning three architectures (gpt2, esm2, mamba) and four
scales (XS/S/M/L), each trained from scratch on UniRef50 under a matched,
Chinchilla-style data budget. The suite is built for controlled comparison —
same data, same tokenizer, one variable at a time.This is a masked-LM pseudo-perplexity over masked positions — a different quantity from the autoregressive models' bits-per-residue. Compareesm2models to each other and on downstream / probing tasks, not via this number againstgpt2/mamba.
| Architecture | esm2 |
| Objective | MLM (masked-language-modeling) |
| Scale rung | XS |
| Parameters | 7.4M (7,422,080) |
| Layers (depth) | 6 |
| Hidden size (d_model) | 320 |
| Attention heads | 20 |
| Max sequence length | 512 |
| Vocabulary | 33-token residue alphabet (ESM-2) |
| norm | layernorm |
| MLP activation | gelu |
| Precision | bf16 |
| Data | UniRef50 release 2026_01 (28-Jan-2026), 60,251,814 sequences; held-out final shard for validation |
| Tokenizer | esm2 — 33-token residue alphabet (shared across the whole suite) |
| Optimizer | Muon (matrices) + AdamW (embeddings/scalars), weight_decay=0.1 |
| Batch size | 524,288 residues/step |
| Optimizer steps | 169 |
| Residues seen | 89.1M (89,064,960) |
| Param/data ratio | 12.0 (Chinchilla-style) |
| Total FLOPs | 5.002e+15 |
| Wall-clock | 0.01 h (1 min) on 4 GPU(s) |
| Seed | 0 (siblings: see below) |
| nanoprot version | 0.5.0 |
esm2 models to each other and on downstream / probing tasks, not via this number against gpt2/mamba.pip install nanoprot), download this repo, and point the
arch-aware loader at the folder — it works for any nanoprot architecture
(gpt2 / esm2 / mamba), reading the embedded config and selecting the right
tokenizer automatically.1from nanoprot.training.checkpoint import load_pretrained
2
3model, cfg, meta, tokenizer = load_pretrained(
4 "path/to/this/repo", device="cpu", return_tokenizer=True,
5)
6model.eval()
7# meta carries the trained-artifact facts (params, FLOPs, val metric, ...)yagizdevre/nanoprot-esm2-XS (seed 0 is the default; siblings on branches seed1,
seed2). This model's sibling seeds:nanoprot-esm2-XS-s0 — final masked-cross-entropy-bits 3.7480nanoprot-esm2-XS-s1 — final masked-cross-entropy-bits 3.7479nanoprot-esm2-XS-s2 — final masked-cross-entropy-bits 3.7470{gpt2, esm2, mamba} x {XS, S, M, L} x {seed 0,1,2}.
See the nanoprot repository for the
complete grid and the scaling-curve comparisons.1@software{nanoprot,
2 author = {Devre, H. Yagiz},
3 title = {nanoprot: a minimal training framework for protein language models},
4 year = {2026},
5 url = {https://github.com/ygzdvr/nanoprot}
6}config.yaml (also embedded in meta_000169.json).
Re-train with:python -m scripts.train --config config.yaml