Views
No views yet
Qwen/Qwen3-4B-Base quantized to FP8 (8-bit weights).Caveat. Needs compute capability >= 8.9 (Ada/Hopper+) to run fast.
| Source | Qwen/Qwen3-4B-Base |
| Scheme | FP8 (8-bit) |
| Format | compressed-tensors |
| Parameters | 4.0B |
| Size on disk | 4.4 GB |
| Compression | 1.82x smaller than the 8.0 GB source |
| Left unquantized | lm_head |
| Quantized on | RTX A6000 |
| Quantized by | Sohailhosseini |
1vllm serve Sohailhosseini/Qwen3-4B-Base-FP8 \
2 --max-model-len 327681from vllm import LLM, SamplingParams
2
3if __name__ == "__main__":
4 llm = LLM("Sohailhosseini/Qwen3-4B-Base-FP8", max_model_len=32768)
5 out = llm.chat(
6 [{"role": "user", "content": "What is quantization? Answer in one sentence."}],
7 SamplingParams(temperature=0.6, max_tokens=512),
8 )
9 print(out[0].outputs[0].text)recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers and hardware are in the table above.