Gemma 4 12B IT - oQ Quantized
Model Description
This repository contains an oMLX oQ quantized version of Google's Gemma 4 12B IT model.
The model has been quantized using oMLX's sensitivity-aware mixed-precision quantization pipeline, which dynamically allocates precision across model components to preserve quality while reducing memory and storage requirements.
Base Model
- Base Model: google/gemma-4-12B-it
- Model Family: Gemma 4
- Quantization Method: oMLX oQ
- Format: MLX
- License: Gemma License
Quantization Information
This model was created using the oMLX oQ quantization pipeline.
oQ uses mixed-precision quantization instead of applying a uniform bit-width across all tensors. More sensitive model components retain higher precision while less sensitive components are compressed more aggressively.
Benefits
- Reduced memory footprint
- Reduced storage requirements
- Improved quality retention compared to uniform quantization
- Optimized for Apple Silicon inference
Intended Uses
This model is suitable for:
- General chat applications
- Coding assistance
- Research and experimentation
- Local AI assistants
- Agent workflows
- Reasoning tasks
- Content generation
Usage
Python
1from mlx_lm import load, generate
2
3model, tokenizer = load("path/to/model")
4
5response = generate(
6 model,
7 tokenizer,
8 prompt="Explain mixed precision quantization.",
9 max_tokens=512,
10)
11
12print(response)
CLI
1mlx_lm.generate \
2 --model path/to/model \
3 --prompt "Hello!"
Hardware Requirements
Hardware requirements depend on:
- Context length
- Runtime implementation
- Quantization parameters
- Concurrent workloads
Apple Silicon systems are recommended for optimal performance.
Limitations
This model inherits the strengths and limitations of the original Gemma 4 12B IT model.
Quantization may introduce:
- Minor reductions in reasoning quality
- Slight output variations compared to full-precision checkpoints
- Reduced accuracy on some specialized tasks
Users should evaluate the model for their specific use cases.
Acknowledgements
Base Model
Google DeepMind — Gemma 4
Quantization
License
This repository contains a quantized derivative of Gemma 4.
Please refer to the original Gemma license and usage terms before deployment.
Disclaimer
This is a community-produced quantized checkpoint and is not an official Google DeepMind release.