Views
No views yet
<|im_start|>system<|im_sep|>You are Phi, a language model trained by Microsoft to help users. Your role as an assistant involves thoroughly exploring questions through a systematic thinking process before providing the final precise and accurate solutions. This requires engaging in a comprehensive cycle of analysis, summarizing, exploration, reassessment, reflection, backtracing, and iteration to develop well-considered thinking process. Please structure your response into two main sections: Thought and Solution using the specified format:<think>{Thought section}</think>{Solution section}. In the Thought section, detail your reasoning process in steps. Each step should include detailed considerations such as analysing questions, summarizing relevant findings, brainstorming new ideas, verifying the accuracy of the current steps, refining any errors, and revisiting previous steps. In the Solution section, based on various attempts, explorations, and reflections from the Thought section, systematically present the final solution that you deem correct. The Solution section should be logical, accurate, and concise and detail necessary steps needed to reach the conclusion. Now, try to solve the following question through the above guidelines:<|im_end|>{system_prompt}<|end|><|user|>{prompt}<|end|><|assistant|>| Filename | Quant type | File Size | Split | Description |
|---|---|---|---|---|
| Phi-4-reasoning-bf16.gguf | bf16 | 29.32GB | false | Full BF16 weights. |
| Phi-4-reasoning-Q8_0.gguf | Q8_0 | 15.58GB | false | Extremely high quality, generally unneeded but max available quant. |
| Phi-4-reasoning-Q6_K_L.gguf | Q6_K_L | 12.28GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, recommended. |
| Phi-4-reasoning-Q6_K.gguf | Q6_K | 12.03GB | false | Very high quality, near perfect, recommended. |
| Phi-4-reasoning-Q5_K_L.gguf | Q5_K_L | 10.92GB | false | Uses Q8_0 for embed and output weights. High quality, recommended. |
| Phi-4-reasoning-Q5_K_M.gguf | Q5_K_M | 10.60GB | false | High quality, recommended. |
| Phi-4-reasoning-Q5_K_S.gguf | Q5_K_S | 10.15GB | false | High quality, recommended. |
| Phi-4-reasoning-Q4_K_L.gguf | Q4_K_L | 9.43GB | false | Uses Q8_0 for embed and output weights. Good quality, recommended. |
| Phi-4-reasoning-Q4_1.gguf | Q4_1 | 9.27GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| Phi-4-reasoning-Q4_K_M.gguf | Q4_K_M | 9.05GB | false | Good quality, default size for most use cases, recommended. |
| Phi-4-reasoning-Q4_K_S.gguf | Q4_K_S | 8.44GB | false | Slightly lower quality with more space savings, recommended. |
| Phi-4-reasoning-Q4_0.gguf | Q4_0 | 8.41GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| Phi-4-reasoning-IQ4_NL.gguf | IQ4_NL | 8.38GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| Phi-4-reasoning-Q3_K_XL.gguf | Q3_K_XL | 8.38GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| Phi-4-reasoning-IQ4_XS.gguf | IQ4_XS | 7.94GB | false | Decent quality, smaller than Q4_K_S with similar performance, recommended. |
| Phi-4-reasoning-Q3_K_L.gguf | Q3_K_L | 7.93GB | false | Lower quality but usable, good for low RAM availability. |
| Phi-4-reasoning-Q3_K_M.gguf | Q3_K_M | 7.36GB | false | Low quality. |
| Phi-4-reasoning-IQ3_M.gguf | IQ3_M | 6.91GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| Phi-4-reasoning-Q3_K_S.gguf | Q3_K_S | 6.50GB | false | Low quality, not recommended. |
| Phi-4-reasoning-IQ3_XS.gguf | IQ3_XS | 6.25GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| Phi-4-reasoning-Q2_K_L.gguf | Q2_K_L | 6.05GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| Phi-4-reasoning-IQ3_XXS.gguf | IQ3_XXS | 5.85GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| Phi-4-reasoning-Q2_K.gguf | Q2_K | 5.55GB | false | Very low quality but surprisingly usable. |
| Phi-4-reasoning-IQ2_M.gguf | IQ2_M | 5.11GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| Phi-4-reasoning-IQ2_S.gguf | IQ2_S | 4.73GB | false | Low quality, uses SOTA techniques to be usable. |
pip install -U "huggingface_hub[cli]"huggingface-cli download bartowski/microsoft_Phi-4-reasoning-GGUF --include "microsoft_Phi-4-reasoning-Q4_K_M.gguf" --local-dir ./huggingface-cli download bartowski/microsoft_Phi-4-reasoning-GGUF --include "microsoft_Phi-4-reasoning-Q8_0/*" --local-dir ./| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
|---|---|---|---|---|---|---|---|
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |