Runs on-device in the TokForge app.
Pre-converted
Josiefied-Qwen3-4B-abliterated-v2 in MNN format for on-device inference with
TokForge.
Josiefied abliterated v2 by Goekdeniz Guelmez — refined 4B Qwen3 with abliterated safety filters. The v2 iteration improves on the original with better uncensoring and instruction following. Great balance of speed and quality for everyday mobile use.
This model is optimized for
TokForge — a free Android app for private, on-device LLM inference.
Pair with the
TokForge Acceleration Pack for speculative decoding. On our test devices, speculative decoding with dense Qwen3 targets measured +34% to +43% faster decode in chat workloads. Results vary by device and workload.
Actual speed varies by device, thermal state, and generation length. Typical ranges for this model size:
This is an MNN conversion of
Josiefied-Qwen3-4B-abliterated-v2 by
Goekdeniz-Guelmez. All credit for the model architecture, training, and fine-tuning goes to the original author(s). This conversion only changes the runtime format for mobile deployment.