Views
No views yet
compressed-tensors format (llm-compressor 0.11.1)k_scale/v_scale baked in — vLLM uses it automatically, no --kv-cache-dtype flaglm_head, embeddings--max-model-len 118288 measured no-OOM (BF16-KV build: ~49K). Arch max 131072neuralmagic/calibration (LLM), max seq 2048| acc | |
|---|---|
| Overall | 78.91% ± 0.33 |
| Social sciences | 88.20% |
| Other | 82.56% |
| Humanities | 75.39% |
| STEM | 71.52% |
1vllm serve amanwalksdownthestreet/Hermes-4-70B-MXFP8-KV8 \
2 --quantization compressed-tensors \
3 --max-model-len 118288 --gpu-memory-utilization 0.95 --max-num-seqs 2