Views
No views yet
mistralai/Devstral-Small-2-24B-Instruct-2512Notes: Keeplm_headin high precision; calibrate on long, domain-relevant sequences.
sudo docker run --runtime nvidia --gpus all -p 8000:8000 --ipc=host vllm/vllm-openai:nightly --model Firworks/Devstral-Small-2-24B-Instruct-2512-nvfp4 --dtype auto --max-model-len 32768 --tool-call-parser mistral --enable-auto-tool-choice