GGUF quantizations of Nanbeige4.2-3B for use with llama.cpp, KoboldCpp, LM Studio, Jan, Open WebUI, Ollama (via GGUF import), and other GGUF-compatible inference engines.
This repository provides a collection of GGUF quantizations of Nanbeige4.2-3B optimized for local inference across a wide range of hardware configurations.
The model was converted from the original Hugging Face weights to GGUF format using the latest available llama.cpp conversion tools. Both traditional K-quants and importance-aware IQ-quants are included to provide a balance between quality, memory usage, and inference speed.
IQ-Quants were generated using importance matrix (imatrix) quantization for improved quality retention at lower bitrates.
Quantizations were produced using the latest available llama.cpp release at build time.
Credits
Base Model
All model weights, architecture, training, and tokenizer credits belong to the original authors of:
Nanbeige/Nanbeige4.2-3B
GGUF Conversion & Quantization
Converted and quantized for the community by:
Nithin Sai Kumar (NANI-Nithin)
Disclaimer
This repository only provides converted and quantized GGUF files. Please refer to the original model repository for:
Training details
Evaluation results
License information
Intended use guidance
Safety considerations
Users must comply with the original model license and usage restrictions.
Support the Original Authors
If you find this model useful, please consider supporting the original creators by starring, downloading, and contributing feedback to the original repository: