Views
No views yet
tencent/UI-Mate-9B quantized to AWQ-W4A16 (4-bit weights).Caveat. Asymmetric; a few older vLLM kernels prefer symmetric W4A16.
| Source | tencent/UI-Mate-9B |
| Scheme | AWQ-W4A16 (4-bit) |
| Format | compressed-tensors |
| Parameters | 9.4B |
| Size on disk | 8.6 GB |
| Compression | 2.19x smaller than the 18.8 GB source |
| Calibration | HuggingFaceH4/ultrachat_200k, 256 samples |
| Left unquantized | lm_head, re:.*visual.*, re:.*vision_tower.*, re:.*vision_model.*, re:.*vision.*, re:.*multi_modal_projector.*, re:.*merger.* |
| Quantized on | A40 |
| Quantized by | Sohailhosseini |
1vllm serve Sohailhosseini/UI-Mate-9B-AWQ-W4A16 \
2 --max-model-len 327681from vllm import LLM, SamplingParams
2
3if __name__ == "__main__":
4 llm = LLM("Sohailhosseini/UI-Mate-9B-AWQ-W4A16", max_model_len=32768)
5 out = llm.chat(
6 [{"role": "user", "content": "What is quantization? Answer in one sentence."}],
7 SamplingParams(temperature=0.6, max_tokens=512),
8 )
9 print(out[0].outputs[0].text)recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers, calibration set and hardware are in the table above.