Views
No views yet
<|system|>
You are a thoughtful and systematic AI assistant built by ServiceNow Language Models (SLAM) lab. Before providing an answer, analyze the problem carefully and present your reasoning step by step. After explaining your thought process, provide the final solution in the following format: [BEGIN FINAL RESPONSE] ... [END FINAL RESPONSE].
{system_prompt}
<|end|>
<|user|>
{prompt}
<|end|>
<|assistant|>
Here are my reasoning steps:| Filename | Quant type | File Size | Split | Description |
|---|---|---|---|---|
| Apriel-Nemotron-15b-Thinker-bf16.gguf | bf16 | 29.96GB | false | Full BF16 weights. |
| Apriel-Nemotron-15b-Thinker-Q8_0.gguf | Q8_0 | 15.92GB | false | Extremely high quality, generally unneeded but max available quant. |
| Apriel-Nemotron-15b-Thinker-Q6_K_L.gguf | Q6_K_L | 12.62GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, recommended. |
| Apriel-Nemotron-15b-Thinker-Q6_K.gguf | Q6_K | 12.29GB | false | Very high quality, near perfect, recommended. |
| Apriel-Nemotron-15b-Thinker-Q5_K_L.gguf | Q5_K_L | 11.07GB | false | Uses Q8_0 for embed and output weights. High quality, recommended. |
| Apriel-Nemotron-15b-Thinker-Q5_K_M.gguf | Q5_K_M | 10.65GB | false | High quality, recommended. |
| Apriel-Nemotron-15b-Thinker-Q5_K_S.gguf | Q5_K_S | 10.39GB | false | High quality, recommended. |
| Apriel-Nemotron-15b-Thinker-Q4_K_L.gguf | Q4_K_L | 9.61GB | false | Uses Q8_0 for embed and output weights. Good quality, recommended. |
| Apriel-Nemotron-15b-Thinker-Q4_1.gguf | Q4_1 | 9.50GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| Apriel-Nemotron-15b-Thinker-Q4_K_M.gguf | Q4_K_M | 9.11GB | false | Good quality, default size for most use cases, recommended. |
| Apriel-Nemotron-15b-Thinker-Q4_K_S.gguf | Q4_K_S | 8.66GB | false | Slightly lower quality with more space savings, recommended. |
| Apriel-Nemotron-15b-Thinker-IQ4_NL.gguf | IQ4_NL | 8.64GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| Apriel-Nemotron-15b-Thinker-Q4_0.gguf | Q4_0 | 8.63GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| Apriel-Nemotron-15b-Thinker-Q3_K_XL.gguf | Q3_K_XL | 8.58GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| Apriel-Nemotron-15b-Thinker-IQ4_XS.gguf | IQ4_XS | 8.20GB | false | Decent quality, smaller than Q4_K_S with similar performance, recommended. |
| Apriel-Nemotron-15b-Thinker-Q3_K_L.gguf | Q3_K_L | 7.99GB | false | Lower quality but usable, good for low RAM availability. |
| Apriel-Nemotron-15b-Thinker-Q3_K_M.gguf | Q3_K_M | 7.40GB | false | Low quality. |
| Apriel-Nemotron-15b-Thinker-IQ3_M.gguf | IQ3_M | 6.94GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| Apriel-Nemotron-15b-Thinker-Q3_K_S.gguf | Q3_K_S | 6.71GB | false | Low quality, not recommended. |
| Apriel-Nemotron-15b-Thinker-Q2_K_L.gguf | Q2_K_L | 6.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| Apriel-Nemotron-15b-Thinker-IQ3_XS.gguf | IQ3_XS | 6.42GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| Apriel-Nemotron-15b-Thinker-IQ3_XXS.gguf | IQ3_XXS | 5.99GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| Apriel-Nemotron-15b-Thinker-Q2_K.gguf | Q2_K | 5.79GB | false | Very low quality but surprisingly usable. |
| Apriel-Nemotron-15b-Thinker-IQ2_M.gguf | IQ2_M | 5.35GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| Apriel-Nemotron-15b-Thinker-IQ2_S.gguf | IQ2_S | 4.98GB | false | Low quality, uses SOTA techniques to be usable. |
| Apriel-Nemotron-15b-Thinker-IQ2_XS.gguf | IQ2_XS | 4.72GB | false | Low quality, uses SOTA techniques to be usable. |
pip install -U "huggingface_hub[cli]"huggingface-cli download bartowski/ServiceNow-AI_Apriel-Nemotron-15b-Thinker-GGUF --include "ServiceNow-AI_Apriel-Nemotron-15b-Thinker-Q4_K_M.gguf" --local-dir ./huggingface-cli download bartowski/ServiceNow-AI_Apriel-Nemotron-15b-Thinker-GGUF --include "ServiceNow-AI_Apriel-Nemotron-15b-Thinker-Q8_0/*" --local-dir ./| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
|---|---|---|---|---|---|---|---|
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |