Views
No views yet
Step3p7ForConditionalGeneration, 196B total / A11B sparse MoE, 288 routed experts top-8)int4_wo_128, file2file), then losslessly repacked to the compressed-tensors format so that vLLM can load it.lm_head, MoE router (moe.gate, router_bias), share_expert, self_attn.g_proj, dense layers 0-2, MTP layers 45-47, and the full vision tower.vllm serve <this-repo> --trust-remote-code --tensor-parallel-size 2 --enable-expert-parallel