Views
No views yet
ik_llama.cpp.ik_llama.cpp Cohere2-MoE support added in PR
#1945.| File | Quant | Size | SHA256 |
|---|---|---|---|
North-Mini-Code-1.0-ik_llama-Q8_0.gguf | Q8_0 | 33,007,688,960 bytes | 2e8139305d30f31ed7a5834c32113f9c6ce5d004bf0ec9008be6da0a20928a50 |
North-Mini-Code-1.0-ik_llama-Q6_K.gguf | Q6_K | 25,522,707,712 bytes | 8661540adc05ccba8cd90e36ca0f29101586a7e201090bc503db1ca11cb4d37d |
North-Mini-Code-1.0-ik_llama-Q4_K_M.gguf | Q4_K_M | 18,991,947,008 bytes | 0dfed0306ef9e0a7887bac556494e4e36b65b116c48976cb0d8283fd48a006cd |
CohereLabs/North-Mini-Code-1.0effaeda477c041c107d5a3d8c599cb5d6c5878efik_llama.cpp main after PR
#1945, or an
equivalent build with cohere2_moe support1e063a6bd Enhance Cohere2-MoE support by modifying tensor handling and configuration logicllama-quantizechat_template.jinjad8366efb9f07c571tokenizer.ggml.pre=cohere2_moegeneral.architecture=cohere2_moetokenizer.ggml.pre=cohere2_moetokenizer.chat_template is embedded and matches the source chat_template.jinjaik_llama.cpp
build containing Cohere2-MoE / North-Mini-Code support.chat_template.jinja.llama-server with the Cohere2-MoE
runtime path.ik_llama.cpp main after PR
#1945, or another
runtime with equivalent cohere2_moe architecture, tensor-loading, tokenizer,
and graph support.general.architecture=cohere2_moe.