Views
No views yet
train_sft split of the HuggingFaceH4/ultrachat_200k dataset. Responses were regenerated by JetBrains/Mellum2-12B-A2.5B-Thinking.| Base Model | JetBrains/Mellum2-12B-A2.5B-Thinking |
| Format | Safetensors |
| License | Apache 2.0 |
| Validation Hardware | Nvidia H100 |
1python3 launch_vllm.py JetBrains/Mellum2-12B-A2.5B-Thinking \
2 --hidden-states-path /tmp \
3 --target-layer-ids 1 6 12 17 23 28 \
4 -- \
5 --max-model-len 16384 \
6 -dp 2 \
7 --max_num_batched_tokens 16384 \
8 --max_num_seqs 4096 \
9 --gpu-memory-utilization 0.9 \
10 --no-enable-chunked-prefill1torchrun \
2 --standalone \
3 --nproc_per_node=2 scripts/train.py \
4 --verifier-name-or-path JetBrains/Mellum2-12B-A2.5B-Thinking \
5 --data-path "prepared_data" \
6 --on-missing generate \
7 --on-generate delete \
8 --scheduler-type cosine \
9 --draft-vocab-size 32000 \
10 --max-anchors 3072 \
11 --target-layer-ids 1 6 12 17 23 \
12 --speculator-type dflash \
13 --num-layers 5 \
14 --logger trackio \
15 --run-name demo_qwen3_9b \
16 --lr 0.0006 \
17 --epochs 10 \
18 --num-workers 16 \
19 --sliding-window-indices 0 1 2 3 41uv pip install -U vllm \
2 --torch-backend=auto \
3 --extra-index-url https://wheels.vllm.ai/nightly1vllm serve JetBrains/Mellum2-12B-A2.5B-Thinking \
2 --speculative-config '{ \
3 "model": "RedHatAI/Mellum2-12B-A2.5B-Thinking-Dflash", \
4 "num_speculative_tokens": 16, \
5 "method": "dflash" \
6 }'| Dataset | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Avg. Length |
|---|---|---|---|---|---|---|---|---|
| HumanEval | 72.5% | 50.9% | 34.3% | 23.4% | 16.4% | 11.4% | 7.6% | 3.17 |
| math_reasoning | 80.9% | 63.3% | 48.5% | 36.8% | 26.7% | 19.1% | 12.9% | 3.88 |
| qa | 64.1% | 38.4% | 23.0% | 13.7% | 8.3% | 5.0% | 2.4% | 2.55 |
| question | 64.9% | 40.1% | 24.3% | 15.1% | 9.3% | 5.8% | 3.4% | 2.63 |
| rag | 64.9% | 39.4% | 24.0% | 13.6% | 7.8% | 4.3% | 2.3% | 2.56 |
| summarization | 58.5% | 31.9% | 17.3% | 9.3% | 5.2% | 3.0% | 1.7% | 2.27 |
| tool_call | 68.8% | 43.1% | 26.5% | 15.5% | 9.1% | 5.4% | 3.1% | 2.71 |
| translation | 58.4% | 33.6% | 18.9% | 10.0% | 5.5% | 2.9% | 1.5% | 2.31 |
| writing | 66.2% | 40.9% | 25.3% | 15.8% | 10.2% | 6.3% | 3.8% | 2.68 |




