Views
No views yet
1# In virtual environment with speculators installed
2python scripts/prepare_data.py \
3 --model google/gemma-4-31B-it \
4 --data ./regenerated_data.jsonl \
5 --output ./output \
6 --seq-length 163841# In (separate) virtual environment with vllm installed
2CUDA_VISIBLE_DEVICES=0,1 vllm_venv/bin/python scripts/launch_vllm.py \
3 google/gemma-4-31B-it \
4 --target-layer-ids 1 17 29 47 58 \
5 --max-model-len 32768 \
6 --max-num-batched-tokens 32768 \
7 --tensor-parallel-size 2 \
8 --no-enable-chunked-prefill1# In virtual environment with speculators installed
2CUDA_VISIBLE_DEVICES=2,3 torchrun \
3 --standalone \
4 --nproc_per_node 2 \
5 scripts/train.py \
6 --verifier-name-or-path google/gemma-4-31B-it \
7 --speculator-type dspark \
8 --from-pretrained <dflash-checkpoint> \
9 --data-path ./output \
10 --on-missing generate \
11 --on-generate delete \
12 --scheduler-type cosine \
13 --draft-vocab-size 32000 \
14 --max-anchors 1024 \
15 --target-layer-ids 1 17 29 47 58 \
16 --num-layers 5 \
17 --lr 1e-4 \
18 --backbone-lr-scale 0.1 \
19 --epochs 3 \
20 --sliding-window 2048 \
21 --markov-rank 256 \
22 --markov-head-type vanilla \
23 --enable-confidence-head \
24 --confidence-head-with-markov \
25 --draft-hidden-act silu| Base Model | google/gemma-4-31B-it |
| Chat Template | google/gemma-4-31B-it (use /chat/completions endpoint) |
| Format | Safetensors |
| License | Apache 2.0 |
| Validation Hardware | Nvidia h100 |
1# Deploy with speculative decoding
2vllm serve google/gemma-4-31B-it -tp 4 --spec-model RedHatAI/gemma-4-31B-it-speculator.dspark --spec-tokens 7 --spec-method dspark| Dataset | Pos 0 | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Avg. Length |
|---|---|---|---|---|---|---|---|---|
| HumanEval | 85.6% | 72.6% | 61.3% | 52.2% | 44.7% | 37.6% | 30.8% | 4.85 |
| math_reasoning | 89.2% | 78.6% | 67.7% | 59.2% | 50.7% | 42.7% | 35.3% | 5.24 |
| qa | 66.4% | 41.8% | 26.0% | 15.8% | 10.1% | 6.5% | 4.4% | 2.71 |
| question | 73.4% | 51.2% | 35.6% | 26.1% | 19.7% | 15.1% | 11.8% | 3.33 |
| rag | 75.1% | 55.1% | 40.5% | 29.4% | 20.9% | 14.0% | 8.5% | 3.44 |
| summarization | 66.9% | 40.3% | 24.7% | 15.0% | 8.7% | 4.4% | 2.6% | 2.63 |
| tool_call | 69.1% | 49.4% | 35.1% | 25.2% | 19.0% | 13.8% | 9.3% | 3.21 |
| translation | 72.3% | 52.0% | 36.6% | 26.5% | 18.8% | 12.6% | 8.1% | 3.27 |
| writing | 73.4% | 51.3% | 35.6% | 25.9% | 19.6% | 15.0% | 11.7% | 3.32 |
