Views
No views yet
<think> reasoning behavior, and license (MIT) are the base model's.compressed-tensors format (llm-compressor 0.12.1)lm_head, embeddings. KV cache: BF16neuralmagic/calibration (LLM), max seq 10241vllm serve <path-or-repo> --quantization compressed-tensors \
2 --max-model-len 49152 --gpu-memory-utilization 0.95<think>…</think> traces per the base model's chat template (--reasoning-parser deepseek_r1 in vLLM separates them).| Benchmark | DeepSeek paper (BF16) | MXFP8 (this repo) |
|---|---|---|
| MMLU-Pro (5-shot, non-thinking, /v1/completions) | not published | 67.50 ±0.41 |
| GPQA-Diamond | 65.2 | pending* |
| AIME24 | 70.0 (multi-sample avg) | pending* |
<think> budget) — do not compare it against thinking-mode numbers.