Nanbeige4.2-3B-Q4_K_M
Q4_K_M GGUF quantization of Nanbeige/Nanbeige4.2-3B.
Details
- Architecture: nanbeige
- Format: GGUF
- Quantization: Q4_K_M
- Effective BPW: 4.93
- Original size: 7953.53 MiB
- Quantized size: 2451.74 MiB
- Runtime: llama.cpp
Usage
Requires a recent llama.cpp version with nanbeige support.
llama-cli -m Nanbeige4.2-3B-Q4_K_M.gguf -ngl 99
Evaluation
Wikitext-2:
- Perplexity: 26.8796 ± 0.2620
- Context: 512
- Batch size: 512
- GPU offload: 99 layers
Base Model
Nanbeige/Nanbeige4.2-3B
No additional training or fine-tuning was performed.