These models are built and maintained on rented GPU compute. If you want to show some appreciation, a follow on X or a coffee helps keep the releases coming.
This is a GPTQ-Pro 4-bit quantization of Jackrong/Qwopus3.6-27B-v2, built to make this excellent Qwopus/Qwen3.6 model practical to run in vLLM with GPTQ-Marlin kernels and long-context inference.
The goal is simple: preserve as much of the original model's character and capability as possible while making it efficient enough for single-GPU RTX 3090-class vLLM deployments.
This is not a new fine-tune. It is a quantized derivative of the original Qwopus3.6-27B-v2 model.
Thanks to Jackrong for the original Qwopus3.6 model, and to groxaxo for GPTQ-Pro and the Qwen3.6 GPTQ-Pro recipe this quantization was aligned with.
Quantization recipe
Setting
Value
Method
GPTQ-Pro / GPTQModel
Bits
4
Group size
128
Symmetric quantization
true
Desc act
false
True sequential
true
Calibration dataset
WikiText-2 raw train
Calibration samples
256
Sequence length
2048
MSE
2.0
Damp percent
0.05
Damp auto increment
0.01
FOEM alpha
0.25
FOEM beta
0.2
Batch size
1
Preserved modules include vision, lm_head, embeddings, and norms.
Validation showed that this artifact preserves MTP-related configuration metadata, but does not include actual mtp.* tensors in model.safetensors.index.json, so this release should be treated as non-MTP for vLLM speculative decoding.
Post-save compatibility patch:
pad_token_id=248055
tokenizer class patched to Qwen2TokenizerFast when needed for vLLM compatibility
Intended serving setup
This checkpoint is intended for text-only vLLM serving on RTX 3090-class hardware.
Serving context for the published vLLM measurements:
The vLLM numbers were collected on an internal llm-residency deployment. The
custom image recipe is not published yet, so this card does not present that
image as a public reproduction target. The stable serving knobs captured from
the run are listed for context.
Field
Value
Nomad job profile
vllm-qwopus36-base
Served model name
qwopus3.6-27b-v2-gptq-pro-foem-4bit-g128-ns256-v2
Critical flags
--dtype float16, --quantization gptq_marlin, --kv-cache-dtype fp8_e5m2, --reasoning-parser qwen3, --tool-call-parser qwen3_coder, --max-model-len 131072; --max-num-batched-tokens was not set explicitly.
Public vLLM Reproducibility
This artifact has a public reproducibility path on the unmodified upstream vLLM OpenAI image:
--enable-sleep-mode was included in the validation command
Startup note: this dense 27B profile can fail the first cold start after torch compile/profiling with a pessimistic KV-cache check. The public validation passed on the second start when reusing persistent vLLM/Nomad-style cache directories such as TORCHINDUCTOR_CACHE_DIR=/data/vllm-qwopus36-base/torch_compile_cache and VLLM_CACHE_ROOT=/data/vllm-qwopus36-base. Treat startup retry plus persistent compile cache as part of the serving recipe.
Reasoning / thinking mode
This model preserves Qwen3-style reasoning behavior. The validation workload below was run with thinking enabled.
MTP / speculative decoding status
This ns256-v2 artifact should be considered text-only and non-MTP for vLLM speculative decoding as published. config.json advertises mtp_num_hidden_layers=1, but the weight index does not contain source mtp.* tensors. Enabling vLLM MTP against this unpatched artifact produced essentially zero accepted draft tokens and poor throughput.
A separate experimental follow-up artifact restores real MTP tensors and quantizes the large MTP linears:
XReyRobert/Qwopus3.6-27B-v2-MTP-GPTQ-Pro-v1
That MTP-GPTQ artifact works and reaches good draft acceptance, but it was still slower than this non-MTP baseline on a single RTX 3090. For practical 100k-131k serving on 1x RTX 3090, this ns256-v2 non-MTP artifact remains the preferred choice.
RTX 3090 validation status
This checkpoint was validated on an RTX 3090 24GB with vLLM, max_model_len=131072, kv_cache_dtype=fp8_e5m2, prefix caching enabled, and thinking enabled.
Observed vLLM multi-turn agent workload metrics:
Metric
Observed value
Notes
Requests observed
15
Multi-turn agent session calls
vLLM request success count
15/15
No vLLM errors observed during the sample
Average prompt size
33,172 tokens
Real multi-turn workload
Average output size
322 tokens
Real generated responses
Average time to first token
5.70s
Prometheus TTFT summary
Average end-to-end request latency
13.07s
Includes prefill, decode, and serving overhead
Average time per output token
0.0230s/token
vLLM TPOT summary
Decode throughput from TPOT
about 43.5 tok/s
Decode-only estimate
Prefix cache hit ratio
83.2% cumulative
vLLM prefix-cache counters
Live 60s prompt throughput
about 1,917 prompt tok/s
Aggregate observed window
Live 60s generation throughput
about 19.1 generated tok/s
Aggregate over full window, including prefill and idle mix
Live 60s prefix-cache hit ratio
78.9%
Delta over the observed window
These are practical multi-turn serving metrics, not a synthetic benchmark. They are useful for RTX 3090-class long-context serving expectations, especially multi-turn usage with prefix caching.
📊 4. Evaluation & Benchmarks
📊 Evaluation & Performance Metrics
24 May update: MMLU-Pro selected subset quality check for the GPTQ-Pro ns256 quantization, plus RTX 3090 vLLM serving metrics.
📚 MMLU-Pro Subset90.57%317 / 350 resolved local run
Evaluation format: This uses the same 350-question MMLU-Pro subset published in the test_data directory of Jackrong/Qwopus3.6-27B-v2: 7 categories, 50 questions per category. This is not a full MMLU-Pro leaderboard run.
Protocol note: The primary 317 / 350 = 90.57% score is rescored with the official MMLU-Pro compute_accuracy.py L2 answer extractor, without random fallback. The generation prompt follows the MMLU-Pro run_gpt4o.py OpenAI-compatible style (Options are:, (A):, and the system instruction asking for The answer is ...); it is not borrowed from the Qwopus model card or test set. The current MMLU-Pro evaluate_from_apiX.py extractor, which selects the last matching answer, scores the same outputs at 320 / 350 = 91.43%. This card reports the conservative L2/no-random-fallback score as the headline.
Model / run
Correct / Total
Accuracy
Δ vs Qwen
Δ vs Qwopus
This GPTQ-Pro ns256 artifact
317 / 350
90.57%
+5.71 pp
+3.14 pp
Qwopus3.6-27B-v2 reference CSV
306 / 350
87.43%
+2.57 pp
baseline
Qwen3.6-27B-v2 reference CSV
297 / 350
84.86%
baseline
-2.57 pp
Category
Qwen3.6-27B
Qwopus3.6-27B-v2
This GPTQ-Pro ns256
Biology
96%
96%
94%
Business
88%
94%
96%
Computer Science
82%
84%
82%
Mathematics
90%
88%
98%
Physics
76%
86%
92%
Chemistry
74%
80%
86%
Health
88%
84%
86%
Summary: On the selected 350-question MMLU-Pro evaluation set, this GPTQ-Pro ns256 artifact reached 90.57% accuracy using the official MMLU-Pro L2 answer extractor without random fallback. This is above the published Qwopus3.6-27B-v2 reference CSV at 87.43% and Qwen3.6-27B-v2 at 84.86%, but the comparison should be treated cautiously because the local result is a resolved mixed run rather than a single-pass uniform evaluation, and it uses the MMLU-Pro run_gpt4o.py prompt style rather than the current local/vLLM prompt template.
Scope note: non-truncated answers were kept, and only length-truncated cases were regenerated with larger output caps, up to max_tokens=16384. Prompts and instructions were not changed. Three pathological cases still ended with finish_reason=length and no parsed answer. For leaderboard-style comparability, the preferred next run is a single-pass full MMLU-Pro evaluation generated from the official GitHub scripts and submitted in the JSON/CSV format expected by the MMLU-Pro HF Space.
⚡ 4.2 RTX 3090 vLLM Runtime Notes
Metric
Observed value
Prompt tokens
509,541
Completion tokens
515,679
Request elapsed sum
11,859.8s
Completion throughput
43.48 tok/s
Paired comparison on the same 350 question IDs is positive but should be treated cautiously because the subset is small and the local result is not a single-pass uniform run. Against Qwopus3.6-27B-v2, this run has 20 local-only correct answers and 9 Qwopus-only correct answers (+3.14 pp, McNemar p≈0.061). Against Qwen3.6-27B-v2, it has 32 local-only correct answers and 12 Qwen-only correct answers (+5.71 pp, McNemar p≈0.0037).
Compatibility notes
This artifact was built and validated for text-only vLLM serving without speculative decoding. Do not enable MTP on this artifact as published; the mtp.* tensors are absent from the weight index. Vision-related modules were not validated for vision use in this release.
Limitations
Experimental quantization.
MTP/speculative decoding is not supported by this published artifact because mtp.* tensors are missing.
Quality has been checked on Jackrong's 350-question MMLU-Pro subset only; this is not a full MMLU-Pro evaluation or an official leaderboard submission.
The subset result uses the official MMLU-Pro answer extractor, but the prompt style is the MMLU-Pro run_gpt4o.py OpenAI-compatible template, not the current local/vLLM template.
RTX 3090 metrics above are observed workload numbers, not a controlled benchmark suite.
Long-context and tool-calling workflows were validated on the described local vLLM/Hermes setup; behavior may vary on other serving stacks, hardware, or generation settings.