Views
No views yet
MODEL_PATH="Qwen/Qwen3.5-397B-A17B"
OUTPUT_DIR="amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2"
cd /path/to/Quark
python3 examples/torch/language_modeling/llm_ptq/quantize_quark.py \
--model_dir "${MODEL_PATH}" \
--quant_scheme mxfp4 \
--layer_quant_scheme '*linear_attn*' ptpc_fp8 \
--layer_quant_scheme '*self_attn*' ptpc_fp8 \
--no_trust_remote_code \
--model_export hf_format \
--skip_evaluation \
--multi_gpu balanced \
--exclude_layers "lm_head" "*mlp.gate" "mtp.*" "model.visual.*" "*shared_expert_gate*" \
--output_dir "${OUTPUT_DIR}"local-chat-completions with the chat template applied, 5-shot, greedy). The baseline is the
original Qwen/Qwen3.5-397B-A17B checkpoint,
evaluated with the identical recipe on SGLang.| Benchmark | Qwen/Qwen3.5-397B-A17B | amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2 (this model) | Recovery |
|---|---|---|---|
| gsm8k (flexible-extract, 5-shot) | 97.71 | 97.24 | 99.55% |
lm-eval recipe.export SGLANG_USE_AITER=1
export SGLANG_USE_AITER_UNIFIED_ATTN=1
export AITER_FLYDSL_FORCE=1
export SGLANG_MAMBA_SSM_DTYPE=bfloat16
export SGLANG_USE_AITER_FP8_PER_TOKEN=1
python3 -m sglang.launch_server \
--model-path amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2 \
--trust-remote-code \
--host 0.0.0.0 \
--port 30000 \
--tensor-parallel-size 2 \
--attention-backend aiter \
--mem-fraction-static 0.8 \
--model-loader-extra-config '{"enable_multithread_load": true}' \
--watchdog-timeout 1200 \
--disable-radix-cache \
--reasoning-parser qwen3 \
--page-size 16 \
--enable-aiter-allreduce-fusion
gsm8k.yaml, whose only change vs. the stock
lm-eval task is a doc_to_text that asks the model to end its response with #### <number> so the
reasoning model's final answer is parsed correctly:lm_eval --model local-chat-completions --apply_chat_template \
--tasks gsm8k.yaml \
--num_fewshot 5 \
--model_args "model=amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2,base_url=http://127.0.0.1:30000/v1/chat/completions,num_concurrent=64,tokenized_requests=False,max_length=16384" \
--gen_kwargs "max_tokens=12288,temperature=0,top_p=1"