This repository provides a quantized GGUF version of the Hermes-4-14B model. The 4-bit and 5-bit quantized variants retains the model’s strengths in advanced reasoning tasks while reducing memory and compute requirements ideal for efficient inference on resource-constrained devices.