This repository contains GGUF quantized versions of Google's Gemma 3 27B pretrained model, optimized for efficient deployment across various hardware configurations.
Quantization Results
Model
Size (GB)
Compression Ratio
Size Reduction
Q8_0
26.7 GB
53%
47%
Q6_K
20.6 GB
41%
59%
Q5_K
17.9 GB
36%
64%
Q4_K
15.4 GB
31%
69%
Q3_K
12.5 GB
25%
75%
Q2_K
9.8 GB
19%
81%
Quality vs Size Trade-offs
Q8_0: Near-lossless quality, minimal degradation compared to F16
Q6_K: Very good quality, slight degradation in some rare cases
Q5_K: Good quality, good balance between size and performance
Q4_K: Decent quality, noticeable degradation but still usable for most tasks
Q3_K: Reduced quality, more significant degradation
Q2_K: Heavily reduced quality, substantial degradation but smallest file size
Recommendations
For maximum quality: Use F16 or Q8_0
For balanced performance: Use Q5_K or Q6_K
For minimum size: Use Q2_K or Q3_K
For most use cases: Q5_K provides a good balance of quality and size
For extreme size constraints: Q2_K provides the smallest file size but with significant quality degradation
Usage with llama.cpp
These models can be used with llama.cpp and its various interfaces. Example:
bash
1# Running with llama-gemma3-cli.exe (adjust paths as needed)2./llama-gemma3-cli --model Google.Gemma-3-27b-pt.q5_k.gguf --ctx-size 4096 --temp 0.7 --prompt "Write a short story about a robot who discovers it has feelings."
License
This model is released under the same Gemma license as the original model.