Views
No views yet
moonshotai/Kimi-K3@9f62e4e9fffbd0a83ddd60e1c209d828994b3569.compressed-tensors, nvfp4-pack-quantizedinput_activations: null)9f62e4e9fffbd0a83ddd60e1c209d828994b3569retroactive-hf-commit-windowfail; never silently requantizeGrEarl/Kimi-K3-NVFP4A16-RequantizedKimiMoE class. The test first loaded and ran the
block through a purpose-built layout path, then rebuilt it and loaded the same
8,072 source tensors through the real K3 model-class chain:
KimiK3ForConditionalGeneration.load_weights -> AutoWeightsLoader ->
KimiLinearForCausalLM.load_weights -> KimiLinearModel.load_weights -> the
generic process_weights_after_loading traversal.20260727T234215Z20260727T220141Z8c22f32993e94e063cb634a03b7fc2f9ff539621model-00049-of-000096.safetensors (17,916,197,160 bytes)vllm/vllm-openai:kimi-k3vllm/vllm-openai@sha256:fb16b180bd9727600067e16fcd6a6de43fb4db1baf4298ef20b4dbdf6bfa5a0e0.1.dev19262+gb6bbf29dd.d202607272.13.0+cu130, CUDA 13.07168 -> 3584 latent down,
896 routed experts with real top-k 16 and 3584 -> 3072 -> 3584, routed
RMSNorm + 3584 -> 7168 up, and 2 shared experts (7168 -> 6144 -> 7168)4.0, linear beta 25.0MARLIN / MarlinExpertsvllm::moe_forward_shared,
_moe_C::moe_wna16_marlin_gemm, _moe_C::grouped_topk,
_C::situ_and_mul, vllm_ir::rms_norm, routing alignment, and MoE sum[2, 896]; selected IDs [2, 16]; 16 unique in-range
experts per token; finite normalized weights[2, 7168], finite and
nonzero; sampled parameter fingerprints matched exactly and outputs were
bit-exact (max_abs_difference=0)$0.341178; all attempts including the
externally canceled run: $3.892933 < $4.5load_weights mapping and generic
post-load traversal on a reduced one-block wrapper. DefaultModelLoader
index/file traversal was not exercised, and the full model was not constructed.lmsysorg/sglang:kimi-k3 image (linux/amd64 digest
sha256:2e8ef3746b2591287db0f37b4470910ab08974bdaf96fc5820bac35f6d3962bc)
found that its compressed-tensors selector supports FLOAT tensor-group NVFP4
only as W4A4 with input activation scales; it has no matching NVFP4A16
FLOAT tensor-group MoE path for input_activations: null. The gate therefore
returned BLOCKED / DO_NOT_RUN_GPU (20260727T221943Z) before weight download
or GPU allocation.DefaultModelLoader index/file traversal, KDA/MLA/vision integration, TP16,
end-to-end token generation, serving quality, or SGLang load/forward/kernel
dispatch.