We first lightly fine-tuned the base model on a diverse set of highly curated data points. We then “palmerized” it by merging models, followed by another light fine-tuning round. Finally, we adjusted Mamba for maximum token speed. With only 90M parameters, this turned the model into a competitive baseline against models above 125M parameters.
Open research, education, hobby use, modification, and redistribution are permitted. Commercial deployment, internal business use, paid products, hosted APIs, and client work require a separate commercial license from appvoid. For commercial licensing:
nosoyhackercodigo@gmail.com
👋 But hey! If you already have a
donation subscription, you can claim access to a free commercial license through the email above. You will keep the rights as long as the subscription is active.