Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3.5-4B-oQ4-fp16-lm – AI Model by RepublicOfKorokke | AlphaNeural AI
You can deploy this model and start earning money today!
RepublicOfKorokke
/
Qwen3.5-4B-oQ4-fp16-lm
like
0
mlx
safetensors
qwen3_5
oq
quantized
Qwen/Qwen3.5-4B
quantized
apache-2.0
4-bit
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
THIS MODEL IS NOT MULTI MODAL
Qwen3.5-4B-oQ4-fp16-lm
This model was quantized using
oQ
(oMLX v0.3.6) mixed-precision quantization.
float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3/M4 and for numerical stability.
Benchmark (on M1 Max)
Model
4k PP
TG
32k PP
TG
Qwen3.5-4B-oQ4
522.7
68.7
345.5
38.9
Qwen3.5-4B-oQ4-fp16
879.6
71.9
466.4
49.6