Runs on-device in the TokForge app.
Pre-converted
DeepSeek R1 Distill Llama 8B in MNN format for on-device inference with
TokForge.
DeepSeek's R1 reasoning capability distilled into a Llama 3.1 8B body. Brings chain-of-thought reasoning to mobile devices. Shows its thinking process step-by-step, making it excellent for math, logic puzzles, coding, and complex analysis. Performance comparable to OpenAI o1 on reasoning tasks.
This model is optimized for
TokForge — a free Android app for private, on-device LLM inference.
Actual speed varies by device, thermal state, and generation length. Typical ranges for this model size:
This is an MNN conversion of
DeepSeek R1 Distill Llama 8B by
DeepSeek. All credit for the model architecture, training, and fine-tuning goes to the original author(s). This conversion only changes the runtime format for mobile deployment.