Views
No views yet
model.safetensors + Eagle3LlamaForCausalLM
config). Serve it as the speculative model:1vllm serve t-tech/T-pro-it-2.1-FP8 \
2 --speculative-config '{"model":"VirVen/T-pro-it-2.1-EAGLE_V3","method":"eagle3","num_speculative_tokens":5}' \
3 --dtype bfloat16 \
4 --tensor-parallel-size 2t-tech/T-pro-it-2.1-FP8 (Qwen3-32B architecture, FP8)