Views
No views yet
| File Name | Size | Recommended Use |
|---|---|---|
Cozum-4B_BF16.gguf | 8.42 GB | Unquantized 16-bit: Maximum precision, but exceeds standard 8GB VRAM limits. Will require system RAM offloading and run slower. |
Cozum-4B_Q8_0.gguf | 4.48 GB | Maximum Quality (8-bit): Near-unquantized performance. Ideal for dedicated GPUs with 8GB+ VRAM, fitting perfectly with room for context. |
Cozum-4B_Q6_K.gguf | 3.46 GB | High Quality (6-bit): Very minimal degradation. Leaves excellent VRAM headroom for large system prompts or extended conversations. |
Cozum-4B_Q5_K_M.gguf | 3.07 GB | The Sweet Spot (5-bit): Excellent balance of quality and size. Highly recommended for general use and coding. |
Cozum-4B_Q4_K_M.gguf | 2.71 GB | Mainstream Standard (4-bit): The best balance of speed and acceptable quality for mid-range hardware. |
Cozum-4B_Q3_K_M.gguf | 2.26 GB | Low-End Hardware (3-bit): Maximum compression for severe memory constraints. |
Cozum-4B_Q2_K_L.gguf | 2.07 GB | Extreme Compression (2-bit): Only use if absolutely necessary; expects noticeable degradation in complex logic and generation quality. |
Cozum-4B_BF16-mmproj.gguf | 676 MB | Multimodal Projection: Required file if utilizing the model's vision/image processing capabilities. |
1<|im_start|>system
2You are a helpful assistant.<|im_end|>
3<|im_start|>user
4Hello!<|im_end|>
5<|im_start|>assistant