A ~50M-parameter
Mamba (selective state-space) causal
language model,
pretrained from scratch on English Wikipedia. Mamba replaces the attention
mechanism of a Transformer with a selective state-space layer, giving linear-time sequence
processing instead of the quadratic cost of self-attention.
Pretrained on the English subset of Wikipedia: over 3 million articles.
1@article{gu2023mamba,
2 title={Mamba: Linear-Time Sequence Modeling with Selective State Spaces},
3 author={Gu, Albert and Dao, Tri},
4 journal={arXiv preprint arXiv:2312.00752},
5 year={2023}
6}