Runs on-device in the TokForge app.
Pre-converted
Josiefied-Qwen3-8B-abliterated-v1 in MNN format for on-device inference with
TokForge.
Josiefied abliterated v1 by Goekdeniz Guelmez — 8B Qwen3 with abliterated safety filters. Excellent quality-to-speed ratio for flagship phones. Runs comfortably on 12GB+ RAM devices with OpenCL GPU acceleration.
This model is optimized for
TokForge — a free Android app for private, on-device LLM inference.
Pair with the
TokForge Acceleration Pack for speculative decoding. On our test devices, speculative decoding with dense Qwen3 targets measured +34% to +43% faster decode in chat workloads. Results vary by device and workload.
Actual speed varies by device, thermal state, and generation length. Typical ranges for this model size:
This is an MNN conversion of
Josiefied-Qwen3-8B-abliterated-v1 by
Goekdeniz-Guelmez. All credit for the model architecture, training, and fine-tuning goes to the original author(s). This conversion only changes the runtime format for mobile deployment.