Views
No views yet
dphn/dolphin-2.9.1-yi-1.5-34bNotes: Keeplm_headin high precision; calibrate on long, domain-relevant sequences.
sudo docker run --runtime nvidia --gpus all -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model Firworks/dolphin-2.9.1-yi-1.5-34b-nvfp4 --dtype auto --max-model-len 8192