This is an experimental quantization of MiniMax M3 to NVFP4 for use on DGX Spark. (Note: This quantization is not DGX Spark only.)
For RTX Pro 6000 Users (or DGX Spark Users who don't want to use sparkrun):
You can run this using the custom sglang container:
MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
M3 is powered by
MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.
We recommend the following inference frameworks (listed alphabetically) to serve the model: