A
TurboQuant 3-bit quantized version of
MiniMax-M2.7, optimized for inference with
turboquant-vllm.
This quantized model is designed to work with the turboquant-vllm inference engine. Please refer to the
turboquant-vllm repository for installation and usage instructions.
This is a 3-bit quantized checkpoint intended for efficient inference. The quantization was applied using the TurboQuant method via the turboquant-vllm project.
This is a third-party quantized version of the original MiniMax-M2.7 model. Please refer to the original model card for base model details and licensing.