I don't want to waste your time reading a 500-word LLM-generated essay on what the model is, when Mistral themselves already provide a good explanation in the original model's README. Go read that instead.
Instead, I'll focus on the important part: why should you use my quantization?
Properly labeled as mistral4 architecture instead of deepseek2. This is a bug in upstream llama.cpp's GGUF conversion code. PR incoming.
Chat template from Leanstral-2603 embedded inside GGUF, no need to specify a template by yourself.
I run these models myself on my Strix Halo box.
You can interrogate me on the Lean Zulip if you find these quants to be malicious, or if you just have suggestions for improvements.