Views
No views yet
llama.cpp with the --no-nextn conversion flag to ensure compatibility across local inference runtimes.| File Name | Quantization Type | File Size | Description |
|---|---|---|---|
hydra-turbo-q4_0.gguf | Q4_0 | 4.95 GB | Legacy 4-bit quantization. Fast execution, low VRAM footprint. |
hydra-turbo-q4_k_m.gguf | Q4_K_M | 5.24 GB | Recommended medium 4-bit quant. Balanced accuracy and memory usage. |
hydra-turbo-q5_k_m.gguf | Q5_K_M | 6.02 GB | High-precision 5-bit quant. Reduced perplexity loss with minimal speed overhead. |
hydra-turbo-q8_0.gguf | Q8_0 | 8.87 GB | Near-lossless 8-bit quantization. Maximum output quality. |
1<|im_start|>system
2You are a helpful assistant.<|im_end|>
3<|im_start|>user
4Your query here<|im_end|>
5<|im_start|>assistant