Runs on-device in the TokForge app.
Pre-converted
Llama 3.2 3B Instruct in MNN format for on-device inference with
TokForge.
Meta's official Llama 3.2 3B Instruct — the compact powerhouse of the Llama family. Designed specifically for edge and mobile deployment. Excellent instruction following in a package that runs on 8GB+ phones. Supports 128K context and 8 languages.
This model is optimized for
TokForge — a free Android app for private, on-device LLM inference.
Actual speed varies by device, thermal state, and generation length. Typical ranges for this model size:
This is an MNN conversion of
Llama 3.2 3B Instruct by
Meta. All credit for the model architecture, training, and fine-tuning goes to the original author(s). This conversion only changes the runtime format for mobile deployment.