Views
No views yet
pack-quantized format. 1.29 TB total (vs 1.56 TB for the original MXFP4
release; a BF16 render of this model would be ~5.6 TB), sized so the full
model fits resident on a single 8x B200 / 8x H200-class node.language_model.model.layers.{1..92}.block_sparse_moe.experts.{0..895}.{w1,w2,w3},
247,296 Linears). Everything else (attention including KDA/MLA, shared
experts, router, latent projections, norms, embeddings, lm_head, vision
tower) ships unchanged in BF16.