Views
No views yet
thinkingmachines/Inkling-Small-NVFP4; the exact matching target checkpoint and tokenizer are required.block_size=15.original_max_position_embeddings=8192, max_position_embeddings=1048576 (matches the target addressable range); config carries both rope_scaling and rope_parameters schemas.mask_token_id=200064.[5, 11, 23, 29, 35].acc_len) at temperature 0 (128 prompts per task, DSPARK block size 7, thinking effort 0.99):| Dataset | acc_len |
|---|---|
| GSM8K | 4.7871 |
| MATH500 | 4.1426 |
| MBPP | 3.4394 |
| HumanEval | 3.3491 |
| MT-Bench | 3.1140 |
| LiveCodeBench | 2.9288 |
| AIME25 | 2.8938 |
| Alpaca | 2.7823 |
| Arena-Hard-v2 | 2.6982 |
| Mean | 3.3484 |
1SGLANG_ENABLE_UNIFIED_RADIX_TREE=1 \
2sglang serve --trust-remote-code \
3 --model-path thinkingmachines/Inkling-Small-NVFP4 --tp 8 \
4 --quantization modelopt_fp4 --attention-backend fa4 --page-size 128 \
5 --fp4-gemm-backend flashinfer_trtllm --moe-runner-backend flashinfer_trtllm_routed \
6 --enable-torch-symm-mem --mamba-radix-cache-strategy extra_buffer \
7 --mem-fraction-static 0.68 --swa-full-tokens-ratio 0.1 --mamba-full-memory-ratio 0.1 \
8 --max-running-requests 68 --reasoning-parser inkling --tool-call-parser inkling \
9 --skip-server-warmup --speculative-algorithm DSPARK \
10 --speculative-draft-model-path skx618/InkLing-Small-DSpark-Preview \
11 --speculative-draft-model-quantization unquant \
12 --speculative-dspark-block-size 7 \
13 --chunked-prefill-size 8192 --cuda-graph-max-bs-prefill 8192 \
14 --disable-flashinfer-autotune --host 0.0.0.0 --port 30000open-perfectblend (temperature 1.0, top-p 0.95, reasoning effort 0.99), 2 epochs, AdamW, global batch 512, gradient clipping 1.0; peak LR 6e-4 (cosine, 4% warmup).max_length=65536; constant LR 5e-4 with a 64-step re-warm.0.1 CE + 0.9 L1 distillation + 1.0 confidence BCE, 512 sampled anchors per sequence, block_size=15, within-block decay gamma 28/3.