MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
MiniMax-M3-MXFP8 is the MXFP8 quantized variant of
MiniMax-M3, a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
M3 is powered by
MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.
We recommend the following inference frameworks (listed alphabetically) to serve the model: