Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
GLM-OCR-oQ8-fp16 – AI Model by RepublicOfKorokke | AlphaNeural AI
You can deploy this model and start earning money today!
RepublicOfKorokke
/
GLM-OCR-oQ8-fp16
like
0
mlx
safetensors
glm_ocr
oq
quantized
image-to-text
zh
en
fr
es
ru
de
ja
ko
zai-org/GLM-OCR
quantized
mit
8-bit
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
GLM-OCR-oQ8-fp16
This model was quantized using
oQ
mixed-precision quantization.
float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3/M4 and for numerical stability.
Benchmark (on M1 Max)
Model Variant
PP (Tokens per second)
TG (Tokens per second)
Original (bf16)
4,684
104.8
oQ8-fp16
3,806
99.0