Views
No views yet
MUL_MAT_ID for the experts) - no dequant-to-f16 fallback.F8E4M3), block-scaled, produced by AMD Quark from BF16.qwen3.6-35b-a3b-quark-fp8-moe-authentic.gguf (38.7 GB).| Metric | Value |
|---|---|
| Perplexity (wikitext, 20 chunks, n_ctx=512), 2-GPU | 6.6270 |
Prefill pp512, 2-GPU | 3197 t/s |
Decode tg128 (single-stream), 2-GPU | 70.2 t/s |
Aggregate decode @ peak concurrency (npl=114) | ~537 t/s (8.2×) |
llama-perplexity on dual R9700 (gfx1201). This is an SSM-hybrid
architecture: use llama-server / llama-perplexity, not llama-bench.llama-batched-bench on dual R9700:Concurrent seqs (npl) | Aggregate decode (tok/s) | Scaling |
|---|---|---|
| 1 | 65.8 | 1.00× |
| 8 | 249.9 | 3.80× |
| 32 | 355.3 | 5.40× |
| 64 | 432.2 | 6.57× |
| 110 | 531.5 | 8.08× |
| 114 | 537.4 | 8.17× ← peak |
| 118 | 535.4 | 8.14× |
npl mod 4 (a batch/ubatch-tiling alignment effect,
reproducible across runs): npl ≡ 2 mod 4 is the favorable alignment (110/114/118
= 531–537 t/s) and ≡ 0 mod 4 is worst (112/116 = 451–490) — so pick a ≡2-mod-4
--parallel value. Same continuous-batching amortization that was vLLM's "server
win," but native and with fp8 (vLLM silently dequantizes fp8 on gfx1201, so its
edge evaporates here). Note: fp8 KV-cache (-ctk/-ctv f8e4m3) is incompatible
with batched decode on this hybrid-SSM MoE (breaks B>1) — batching is f16-KV only.
For MoE, reach for concurrency (this section); for single-user latency, MTP below.1# -ngl 999 lets llama.cpp see and tensor-split across both R9700s
2llama-server -m qwen3.6-35b-a3b-quark-fp8-moe-authentic.gguf -ngl 999 --host 0.0.0.0 --port 13305
3# it has an MTP head -> self-speculative decode for faster single-stream latency:
4llama-server -m qwen3.6-35b-a3b-quark-fp8-moe-authentic.gguf -ngl 999 --spec-type draft-mtp
5
6# multi-user: continuous batching for concurrent serving (the MoE's strength)
7# (f16-KV only — fp8-KV breaks batched decode on this hybrid-SSM MoE)
8llama-server -m qwen3.6-35b-a3b-quark-fp8-moe-authentic.gguf -ngl 999 \
9 --cont-batching --parallel 114 # peaks ~537 tok/s aggregate (npl=114, a ≡2-mod-4 value)
10
11curl -s http://localhost:13305/v1/chat/completions -H 'Content-Type: application/json' \
12 -d '{"messages":[{"role":"user","content":"What do you call a dried grape? Answer in one word."}],"max_tokens":16}'
13# expect: raisin1podman run -d --rm --runtime crun --name lemonade \
2 --device /dev/kfd --device /dev/dri \
3 --group-add keep-groups --security-opt seccomp=unconfined \
4 -v /path/to/quacken-35b:/models:ro \
5 -e MODEL=/models/qwen3.6-35b-a3b-quark-fp8-moe-authentic.gguf -e MODEL_NAME=Quacken-35B-A3B-FP8 \
6 -p 13305:13305 \
7 ghcr.io/the-monk/the-rock8:rdna4-tr713 serve
8# 35B needs both GPUs - do NOT pin HIP_VISIBLE_DEVICES to a single card--runtime crun is required for GPU):
ghcr.io/the-monk/the-rock8:rdna4-tr713 - docker.io/gorilla4x/the-rock8:rdna4-tr713 - quay.io/the-monk/the-rock8:rdna4-tr713
(images may not be pushed to every registry yet).