Views
No views yet
llm-compressor's latest changes, quantized on a GH200, works well for me with vLLM's main branch on my RTX 3090Ti as of 2025-07-01.MistralTokenizer.config.json to be compatible with this approach, and I'm about to push the tekken.json tokenizer. With that, if you build that branch, you should be able
to run this checkpoint with MistralTokenizer and get tool calling.