MicroBananaMind-v1 is a very small causal language model trained from scratch on FineWeb-Edu, FineMath, and Cosmopedia-v2.
The model has 902,272 parameters and uses a custom 1536-token byte-level BPE tokenizer with digit-aware tokenization
It is our smallest model ever that is not just a TinyStories model.