Views
No views yet
mamba architecture) — 151.7M (151,749,504) parameters, trained on 1.82B (1,820,994,048) UniRef50 residues.nanoprot-mamba-M is part of the nanoprot suite: a Pythia-style matrix of protein
language models spanning three architectures (gpt2, esm2, mamba) and four
scales (XS/S/M/L), each trained from scratch on UniRef50 under a matched,
Chinchilla-style data budget. The suite is built for controlled comparison —
same data, same tokenizer, one variable at a time.Directly comparable to other autoregressive nanoprot models (gpt2,mamba) — same 33-token vocabulary, same AR objective, same data budget. Not comparable to theesm2(masked-LM) models, whose metric is a different quantity.
| Architecture | mamba |
| Objective | AR (causal-language-modeling) |
| Scale rung | M |
| Parameters | 151.7M (151,749,504) |
| Layers (depth) | 18 |
| Hidden size (d_model) | 1152 |
| Attention heads | 1 |
| Max sequence length | 512 |
| Vocabulary | 33-token residue alphabet (ESM-2) |
| d_state | 16 |
| d_conv | 4 |
| expand | 2 |
| dt_rank | 72 |
| Precision | bf16 |
| Data | UniRef50 release 2026_01 (28-Jan-2026), 60,251,814 sequences; held-out final shard for validation |
| Tokenizer | esm2 — 33-token residue alphabet (shared across the whole suite) |
| Optimizer | Muon (matrices) + AdamW (embeddings/scalars), weight_decay=0.1 |
| Batch size | 524,288 residues/step |
| Optimizer steps | 3,473 |
| Residues seen | 1.82B (1,820,994,048) |
| Param/data ratio | 12.0 (Chinchilla-style) |
| Total FLOPs | 5.361e+18 |
| Wall-clock | 1.52 h (91 min) on 4 GPU(s) |
| Seed | 0 (siblings: see below) |
| nanoprot version | 0.5.0 |
gpt2, mamba) — same 33-token vocabulary, same AR objective, same data budget. Not comparable to the esm2 (masked-LM) models, whose metric is a different quantity.pip install nanoprot), download this repo, and point the
arch-aware loader at the folder — it works for any nanoprot architecture
(gpt2 / esm2 / mamba), reading the embedded config and selecting the right
tokenizer automatically.1from nanoprot.training.checkpoint import load_pretrained
2
3model, cfg, meta, tokenizer = load_pretrained(
4 "path/to/this/repo", device="cpu", return_tokenizer=True,
5)
6model.eval()
7# meta carries the trained-artifact facts (params, FLOPs, val metric, ...)yagizdevre/nanoprot-mamba-M (seed 0 is the default; siblings on branches seed1,
seed2). This model's sibling seeds:nanoprot-mamba-M-s0 — final bits-per-residue 3.7841nanoprot-mamba-M-s1 — final bits-per-residue 3.7865nanoprot-mamba-M-s2 — final bits-per-residue 3.7853{gpt2, esm2, mamba} x {XS, S, M, L} x {seed 0,1,2}.
See the nanoprot repository for the
complete grid and the scaling-curve comparisons.1@software{nanoprot,
2 author = {Devre, H. Yagiz},
3 title = {nanoprot: a minimal training framework for protein language models},
4 year = {2026},
5 url = {https://github.com/ygzdvr/nanoprot}
6}config.yaml (also embedded in meta_003473.json).
Re-train with:python -m scripts.train --config config.yaml