Views
No views yet
OpensourceWTF/Kimi-K3-Q2_K-t158-MTPLX-streaming.[!IMPORTANT] This is not Moonshot AI's original Kimi K3 checkpoint, and it is not a llama.cpp/GGUF model. The 92 routed MoE layers were requantized fromGrEarl/Kimi-K3-GGUFQ2_K to the MTPLXt158codec. No downstream quality evaluation has been completed.
model-00001-of-00188.safetensors through
model-00096-of-00188.safetensors: resident model tensorsmodel-00097-of-00188.safetensors through
model-00188-of-00188.safetensors: one routed expert layer per shardmodel.safetensors.index.json: all 3,180 tensor-to-shard mappingsmtplx_t158.json: the packed-expert serialization contractgate_proj.packed, gate_proj.scales, up_proj.packed, up_proj.scales,
down_proj.packed, and down_proj.scales. Expert ID is the leading dimension.
The packed weights use Safetensors U8, and the BF16 scale bit patterns use
U16..bin banks, streaming manifests, conversion receipts,
or build journals in this repository.GrEarl/Kimi-K3-GGUF@0169245d3ea1473a3f9f03bca821d855df5fb2a3moonshotai/Kimi-K3@9f62e4e9fffbd0a83ddd60e1c209d828994b3569t158, group size 64, 1.875 physical bits per weight5dc510958199ad6b9d5701e47dd96bff14c41fbec1531a7a93ed16dc0176bcc1U8 and scale-bit U16 tensors, but the
expert tensors require an MTPLX t158-aware loader. A generic Transformers
from_pretrained call does not know how to execute this packed expert layout.0.903718 against the
decoded Q2_K values for one selected expert. That is a narrow conversion
diagnostic, not a model-quality score or benchmark.Kimi K3 License.
See the
moonshotai/Kimi-K3
model card for the original architecture, evaluation, and usage documentation.