Views
No views yet
python convert_hf_to_gguf.py /home/admin/models/unsloth/Kimi-K2-Instruct-BF16 --outfile /home/admin/models/unsloth/Kimi-K2-Instruct-BF16/ggml-model-f16.gguf --outtype f16llama-quantize /home/admin/models/unsloth/Kimi-K2-Instruct-BF16/ggml-model-f16.gguf /home/admin/models/unsloth/Kimi-K2-Instruct-BF16/ggml-model-Q2_K.gguf Q2_Kllama-cli -m /home/admin/models/unsloth/Kimi-K2-Instruct-BF16/ggml-model-Q2_K.gguf -n 2048-- #define LLAMA_MAX_EXPERTS 256 // DeepSeekV3
++ #define LLAMA_MAX_EXPERTS 384 // Kimi-K2-Instructollama run huihui_ai/huihui_ai/kimi-k2:1026b-Q2_Knum_gpu inside the model is 1, which means it defaults to loading one layer. All others will be loaded into CPU memory. You can modify num_gpu according to your GPU configuration./set parameter num_gpu 2/set parameter num_thread 32/set parameter num_ctx 4096 bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge