Views
No views yet
UraionLabs/Uraion-Agent-Steer, converted with
llama.cpp.model.layers.*.hres.*. Current GGUF / llama.cpp runtimes do not implement the
nonlinear H-Res residual adapter path, so these GGUF files omit the unsupported
H-Res tensors during conversion.UraionLabs/Uraion-Agent-Steer.| Quant | File | Size | Suggested use |
|---|---|---|---|
| Q2_K | Uraion-Agent-Steer-Q2_K.gguf | 2.81 GiB | Smallest fallback, lowest quality |
| Q3_K_M | Uraion-Agent-Steer-Q3_K_M.gguf | 3.55 GiB | Low-memory fallback |
| Q4_K_M | Uraion-Agent-Steer-Q4_K_M.gguf | 4.36 GiB | Recommended default for 8 GB GPUs |
| Q5_K_M | Uraion-Agent-Steer-Q5_K_M.gguf | 5.07 GiB | Better quality if memory allows |
| Q6_K | Uraion-Agent-Steer-Q6_K.gguf | 5.82 GiB | High quality local inference |
| Q8_0 | Uraion-Agent-Steer-Q8_0.gguf | 7.54 GiB | Largest compatibility quant |
1llama-cli -hf UraionLabs/Uraion-Agent-Steer-GGUF:Q4_K_M \
2 -p "What is tool calling?"ollama run hf.co/UraionLabs/Uraion-Agent-Steer-GGUF:Q4_K_Mllama.cpp
quantization types: Q2_K, Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, and Q8_0.